All essaysAssessment

A score without a timestamp is a guess

A 1–10 interview rating cannot survive a debrief. Binary checks, cited to the tape, can — and the human still decides.

A race car in motion on a wet track at night

Three interviewers walk into a debrief with three numbers. 7. 8. 6. Someone says “strong communicator.” Someone else says “not quite senior.” Nobody can point to a sentence. Twenty minutes later the panel is arguing about memory, not about the candidate.

That is not a close call. That is an unscored interview wearing a number.

Hireflo’s bar is simpler than a rubric PDF: if you cannot replay the moment, you did not score it.

The 1–10 is a round of applause

A single rating feels like a decision because it is a number. It is not. It compresses an hour of speech into a vibe, then asks three people to average their vibes.

The collapse is predictable:

  • The same “7” means “hire, with nerves” to one interviewer and “pass, but I liked them” to another.
  • “Culture fit” and “communication” soak up everything that was never written down.
  • The loudest memory in the room wins. The quiet evidence does not get a vote.

You can train interviewers. You can add a second round. You cannot unsay a score that was never attached to a clip. The debrief will keep litigating who remembers what.

A score that cannot be replayed is not a score. It is a mood with a digit.

Binary checks, cited to the tape

Hireflo does not ask a model for a 1–10. It asks a different question, per criterion: did this happen in the candidate’s speech, or not?

The loop is short on purpose.

  1. The job description becomes a list of yes/no checks — not “strong system design,” but “named a failure mode and how they would detect it.”
  2. The room runs against that list. Follow-ups are allowed. Filling silence for the candidate is not.
  3. Afterward, each check is pass, miss, or insufficient evidence. A pass is refused unless the quote appears in a candidate turn.
  4. The report shows the quote and the timestamp. You can disagree with the criterion. You cannot disagree with the tape.

That last line is the product. Language models extract the checks, run the interview, and gate the credit. They do not get to launder a hunch into a number.

Gut score Cited score
“Solid 8 on communication.” Pass: explained the tradeoff in their own words at 12:04.
“Seemed senior.” Miss: never named who owned the incident, or how they would page.
“Not sure they go deep.” Insufficient evidence: the question was asked; the answer was a slogan.

Insufficient evidence is not a polite miss. It means the tape does not support a call, so the call is not made. That is stricter than a slider, and kinder than inventing a 6.

What the model is allowed to do

During the session, a speech model converses from the rubric and the question plan you configured. After the session, an assessment model scores each criterion and must quote the span that supports the call.

It does not look at a face. It does not score clothing, the room, or the background. It does not treat accent, pitch, or “confidence” as a signal. Those are easy for software to fake-measure and impossible to defend in a debrief. We do not collect them as evidence because they are not evidence of the work.

It also does not output a single gut score as the decision, and it does not auto-advance or auto-reject anyone. Hireflo recommends. Your team hires. That is policy, not a footnote.

If this sounds slower than a score from 1 to 10, it is — by about the length of a replay. The debrief gets that time back because the argument has a source.

The human still has to sit in the chair

Cited scoring is not a substitute for a hiring manager. It is what you owe the manager so they can do the job.

You can override a criterion. You can throw out a question that was badly written. You can decide that a miss on a nice-to-have is not a miss on the role. What you cannot do is hide behind “the AI said 7.4.” There is no 7.4. There is a list of checks, a set of clips, and a person who will put their name on the outcome.

That is the point of our explainability statement and the Trust Center: the chain is inspectable. Inputs (job, rubric, transcript) → evidence spans → criterion judgments → a report the panel can replay. Nothing in that chain is a face, a voiceprint, or an accent.

If we cannot show the moment, we do not award the criterion. If we cannot stand behind that rule in public, we should not be in the room.

Try it

Run a session against a real job, then open the report and click a score. If the clip is not there, the score should not be there either.

Start a free mock · Check a hiring flow · Read how a score is made