aifluent research

The lab behind the evidence standard.

aifluent research defines what AI-native means and how to measure it: the judgment people show when they work with AI.

1 construct4 dimensionsverify is the spine
01 · The construct

What AI-native means, and doesn’t.

AI-native means judgment with AI on real work, not tool trivia. The lab’s construct rests on four dimensions: use, verify, decide, communicate.

Verify is the spineWeighted heaviest, about 25%: catching what AI gets wrong is the hardest signal to fake. As AI takes over drafting, knowledge work itself shifts toward verification (Lee et al., CHI 2025, 319 knowledge workers).
The jagged frontierIn a 758-consultant experiment, AI raised quality 40% inside its frontier and made people 19 points worse outside it (Dell’Acqua et al., Organization Science). Knowing which side you’re on is the skill.
Judgment, not literacyHuman + AI is not automatically better: people who trust AI more think about its output less, and teams that accept suggestions uncritically underperform the AI alone (Buçinca et al., CSCW). The lab measures the judgment that closes that gap.
Everyone who decidesThe construct covers every role that makes decisions with AI: PM, analyst, chief of staff, HR, ops, engineering.

The full mechanics, end to end: How it works →

02 · Item design

Traps, gates, and unique variants.

The lab’s item-design program builds cases where a careless run fails loudly and a careful run can win from the packet alone: salience asymmetry, by construction. The method stands on two established lineages: evidence-centered design (Mislevy et al., 2003) and behaviorally anchored rating scales (Smith & Kendall, 1963).

Adversarial gatesBait, solvability, and hazard checks run against every case before it ships. Fail a gate and the case regenerates.
RecalibrationAs frontier models improve, cases recalibrate: difficulty stays in the task, never the tool.
Radicals, incidentalsRubric, trap patterns, and difficulty hold fixed across candidates. Names, numbers, and where the detail hides vary per run.
03 · Validity

Science that survives scrutiny.

Every scoring claim is built to stand in front of a regulator, a court, and the candidate it describes.

Work samples hold upThe current meta-analytic consensus puts work samples among the most valid, job-related predictors of performance, alongside structured interviews (Sackett et al., 2022, Journal of Applied Psychology, revising Schmidt & Hunter, 1998).
DefensibilityRubrics with weights and L1 to L4 behavioral anchors (four defined performance levels), aligned to the SIOP Principles (2018) and the Uniform Guidelines (29 CFR 1607), built for NYC LL144 bias audits.
EU AI Act postureHiring AI is high-risk under Annex III (Reg. 2024/1689); Article 14 requires effective human oversight. Built for that scrutiny: human oversight by design, documented end to end.
Scored from the traceThe signal comes from the work itself: decisions, verifications, and the trace behind them.
04 · Why now

The cheating crisis is a design problem.

Gartner projects that by 2028, 1 in 4 candidate profiles worldwide will be fake, and its 2025 survey found only 26% of candidates trust AI to evaluate them fairly. Vendor analyses of live interviews put AI-use flag rates near 38% (Fabric).

The lab’s position: don’t police the tool. Redesign the assessment so cheating is irrelevant, and full-strength AI use is the point.

AI-assisted isn’t AI-native. The difference is measurable.

05 · The evidence

What the benchmark stands on.

Each source labeled for what it is: peer-reviewed, regulation, or industry data.

Organization Science. 758 consultants: +40% quality inside AI's frontier, 19 points worse outside it. Judgment about the frontier is the differentiator.
319 knowledge workers: the more people trust AI, the less they think critically about it. The work shifts toward verification: the spine of the construct.
Journal of Applied Psychology. The current validity consensus: work samples and structured interviews sit in the top cluster of job-related predictors.
Use, verify, decide, communicate: the lab’s taxonomy for AI judgment, scored from a work sample rather than self-reported. Grounded in Anthropic’s AI Fluency framework (Dakan & Feller): delegation, description, discernment, diligence.
EU AI Act · NYC LL144regulation
Reg. 2024/1689 Annex III + Art. 14 (human oversight for hiring AI); NYC's bias-audit law, enforced since 2023. The scrutiny the method is built for.
Gartner, 2025 ↗industry data
1 in 4 candidate profiles projected fake by 2028, and only 26% of candidates trust AI to evaluate them fairly. The industry’s own answer is fraud-resistant, AI-inclusive assessment.