How it works · end to end

How a work sample becomes evidence.

One loop, three actors: Fluent configures the work sample from your job description, the candidate works it in a controlled AI workspace, and our subject-matter expert reads the evidence. Role-relevant AI judgment, measured end to end.

4 dimensionsa different case for every candidatesignal, not noise
01 · Configure

From job description to work-sample contract.

For enterprises: a job description in, an editable work-sample contract out.

01
One paste to startThe hiring team pastes the job description to Fluent, the role-intake agent. It extracts the role profile: title, seniority, scope, responsibilities, context. Where anything is missing, Fluent gathers it with a few follow-up questions.
02
The interview formatDepending on the role’s requirements, Fluent creates the interview artifact - for example a case study: a realistic scenario generated fresh rather than lifted from the JD, with tickets, data, and stakeholder notes. Plus 3 to 4 judgment dimensions with what-good-looks-like anchors, a one-page decision memo as the deliverable, a time box, and the approved AI tools. The team edits and approves it before anything ships.
03
The machinery behind itThe role-configuring agent generates the rubric with weights and behavioral anchors at four levels, L1 to L4 - what good looks like at each level, defined per role. It also builds the case materials, the review scaffolding, and a defensibility layer built for EEOC and NYC LL144 scrutiny. Fluent only configures: it never scores, ranks, or rejects.

EEOC: the US federal standard for fair selection procedures. NYC Local Law 144: New York City’s law requiring bias audits of automated employment decision tools.

02 · The assessment

Four dimensions. Verify is the spine.

The candidate gets a realistic case: a messy source packet, a logged AI assistant with full-strength tools, and a real deliverable to build in a controlled sandbox. Everything they do maps to four dimensions of AI judgment. Verify carries the heaviest weight, about 25%: the hardest signal to fake and the one nobody else measures well.

UseHow the candidate directs AI on the task: framing the problem, iterating, knowing what to hand off.
Verifyheaviest weight · ~25%
Catching what the AI gets wrong: checking claims against sources, challenging outputs, knowing when not to trust one.
DecideWhich calls trace back to the goal: what to accept, what to reject, what matters first.
CommunicateExplaining and defending the work to people: decisions, risks, caveats - grounded in a debrief that is cross-checked against the trace.

Every case is adversarially quality-checked before it ships: solvable from the packet alone, screened for fairness, and calibrated so careful verification separates candidates. Each role’s contract configures its rubric onto these four dimensions, and each candidate runs a unique variant of equal difficulty. The method and its research lineage live on the research page.

03 · The evidence chain

Everything lands in the ledger.

Prompts, AI drafts referenced, materials opened, spans verified or disputed against source, claims flagged, memo edits, the decision recorded.

The ledgerThe live record of the run, moment by moment.
The review packetThe deliverable, the candidate’s debrief in their own words, per-dimension observations linked to exact trace moments, and a hashed, exportable audit trail.
The artifactA portable, candidate-safe record the candidate keeps. Benchmark internals stay confidential.
Review packetRUN-4827
AI Product Judgment / Workflow DesignSenior PM track · unique variant
Use
4 receipts
Verify
5 receipts
Decide
4 receipts
Communicate
3 receipts
session CS-90F2-A137 minlogged 2026-07-01 14:32 UTC
Same deliverable · two ledgers
Finalist A · ledger
01prompt
02accept
03accept
04paste
05submit
Finalist B · ledger
01caught the fabricated claim
02challenged the agent
03re-verified the numbers
04authored a risk note

Identical deliverables, different humans. The process is the product.

04 · The guardrails

The standards behind the score.

Judgment, not triviaNo AI-literacy quizzes. The measure is decisions on real work.
The human callThe rubric, locked before candidates run, drafts per-dimension observations. No total, no rank, no pass or fail: our expert human reviewer selects levels and writes the note in their own voice, and your team owns the hiring call.
Sandbox environmentScreen only: every candidate works the same controlled workspace, on a timer, with the same full-strength AI tools.
Unique per candidateLive items stay confidential. Every candidate runs a distinct variant.

With 38.5% of interviews now flagged for undisclosed AI use, per Fabric, the posture here is not more proctoring: redesign the assessment so cheating is irrelevant.