Technology hiring · 6 MIN READ
By Alivio Search Partners ·
Sources checked September 21, 2026. Examples and templates are illustrative, not client results.
A technical interview scorecard records the competencies a role requires, the evidence collected during assessment, and the criteria used to interpret that evidence. Its purpose is to make a hiring decision explainable. It should help interviewers distinguish an observed skill, an unanswered question, and a personal impression.
Start with the work the person will do. A scorecard for a backend engineer maintaining a production service should not automatically be reused for a data analyst or a first engineering manager. The template and exercises below are illustrative; they are not a validated assessment instrument or a record of actual candidates.
Define the role before designing the interview
Write three or four outcomes the person should contribute to during an agreed period. For a backend role, those outcomes might concern reliable changes to an existing service, diagnosing failures, and communicating trade-offs. Confirm them with the manager and the people who understand the work.
Then identify the competencies you need evidence for. Keep the list short enough to assess meaningfully. If your process contains twelve criteria but only one short conversation, you may produce more unsupported ratings than useful evidence.
The U.S. Office of Personnel Management’s structured-interview guidance describes using questions about past behavior or hypothetical situations to assess job-related competencies. That principle supports a consistent interview structure. Your team still needs to define the role-specific content and evaluate whether its assessment is appropriate.
An illustrative backend-engineer scorecard
| Competency | Example evidence to collect | Question that remains if evidence is missing |
|---|---|---|
| Problem diagnosis | Candidate separates observations from hypotheses and proposes a useful next check | Can they explain how a check would change the diagnosis? |
| Change safety | Candidate identifies tests, rollout considerations, and a way to detect a regression | Which failure would their proposed test actually catch? |
| Data and interface reasoning | Candidate explains relevant data assumptions and API behavior | What happens when an input or downstream dependency behaves unexpectedly? |
| Trade-off communication | Candidate describes options, constraints, and why one approach fits the situation | Can another engineer follow the reasoning without guessing? |
| Collaboration | Candidate explains a concrete example of incorporating feedback or resolving a disagreement | What did they personally do, and what changed? |
Use this table as a starting point for discussion, not a requirement to run five separate interviews. Several competencies may be assessed in one well-designed exercise. Avoid scoring the same anecdote repeatedly as if it were independent evidence.
Use ratings that describe evidence
A numerical average can hide an important gap. Start with language the team can interpret consistently:
- Not assessed: the process did not collect enough relevant evidence.
- Below the agreed requirement: observed evidence did not meet the defined criterion.
- Meets the agreed requirement: the evidence supports the criterion for this role.
- Exceeds the agreed requirement: the evidence goes beyond that criterion in a way relevant to the work.
For each rating, record a brief observation and its source. “Proposed checking request traces before changing a timeout” is an observation. “Great instincts” is an interpretation without enough supporting detail. A missing assessment should not silently become a low score.
If you use numbers, define them before interviewing. Do not add weights after seeing a preferred candidate’s results. Discuss which requirements are essential and which gaps can reasonably be addressed through onboarding.
Design a small, realistic exercise
For the illustrative backend role, you could present a fictional service with an intermittent failure, a short incident description, and a limited set of logs. Ask the candidate what they would check next and what information they need. The exercise should not require access to production systems, customer data, or confidential code.
Tell candidates what the assessment covers, its expected duration, and the tools they may use. If AI assistance is allowed, say how its use will be evaluated. For example, the discussion may focus on whether the candidate can explain and verify a suggested change. If it is not allowed for a specific assessment, communicate that beforehand and keep the exercise aligned with the skill being assessed.
Provide a way to request an accommodation through the organization’s established process. Consistency means applying the same job-related criteria; it does not require ignoring an individual’s access needs.
Agree on follow-up questions in advance
Useful follow-ups probe the reasoning behind an answer. Examples for the fictional service exercise include:
- What would you expect to see if your hypothesis were correct?
- Which alternative explanation would you rule out next?
- How would you test the change before releasing it?
- What would make you stop or reverse the rollout?
- How would you explain the decision to another team?
Avoid introducing unrelated puzzles because an interviewer has spare time. If a new question reveals that the brief is missing a competency, record that as a process issue. Decide whether and how it should be assessed consistently for the candidates still under consideration.
Record evidence before the debrief
Ask interviewers to complete their notes independently before discussing an overall recommendation. This gives the group specific evidence to examine rather than a first speaker’s verdict to react to.
Organize the debrief around three categories: criteria supported by evidence, concerns supported by evidence, and unanswered questions. When ratings differ, look at what each interviewer observed and whether they applied the same definition. A disagreement may reveal a different interpretation of the role rather than a difference in candidate ability.
The hiring manager should explain the final decision against the agreed requirements. The scorecard supports judgment; it does not make that judgment automatically.
A hypothetical example of better notes
An unhelpful note says: “Weak debugging; probably too junior.” A more useful note says: “During the fictional timeout exercise, the candidate proposed increasing the timeout without first checking request duration or downstream errors. When asked what evidence would distinguish those causes, they did not identify a check.”
The second note does not diagnose the candidate’s overall capability. It records what happened in one assessment and the question it leaves open. Another part of the process may provide relevant evidence. Keep the scope of the conclusion proportional to the information collected.
Similarly, a polished answer is not proof that the candidate performed the work they describe. Ask about their own contribution, constraints, and what they would change. Record those details without requesting confidential information from a previous employer.
Review the scorecard after hiring
Once a person has started, compare the role assumptions with the work they actually encounter. Were the assessed competencies relevant? Did the process overlook a recurring responsibility? Were some questions confusing or redundant?
Treat that review as a way to improve the process. Avoid claiming that a small number of hires proves the assessment predicts performance. Use documented observations, feedback from the team, and appropriate expertise when making broader changes.
Download our candidate-scorecard template and adapt it to the role. Alivio’s technology recruiting service can be scoped around your hiring priorities; share a brief to discuss the evidence your team needs.
Frequently asked questions
Should every candidate receive identical questions?
Use a consistent core structure and criteria for the same role, with appropriate follow-ups to understand an answer. Decide how to handle accommodations and equivalent evidence through your organization’s process.
Should we average every interviewer’s score?
An average can conceal a missing assessment or disagreement about an essential requirement. Review evidence and unresolved questions before using any overall numerical summary.
Can AI score the interviews for us?
Do not treat an automated score as a hiring decision. If you use AI to organize notes, review the output for accuracy and keep the original evidence, role criteria, and accountable human decision-maker visible.
