Work simulations for hiring
Truly turns your job description into a job-specific work simulation. Candidates do the actual work on their own machine, with their own tools — and AI evaluators turn the full session into evidence you can hire on.
100+ candidates assessed · 3+ startups hiring with Truly today
Evaluation
Strong hire signal. Plans before coding, uses AI to reason — not to outsource — and communicates risk early.
The problem
AI changed how work gets done. Screening didn't keep up — so companies still gamble on candidates they've never actually seen work.
Generic, sandboxed challenges pose problems a candidate will never see again at work. Passing one predicts trivia recall, not on-the-job performance.
Behavioral rounds measure preparation, not character. Culture fit — the thing early teams live or die by — never actually gets tested.
Human review is slow and inconsistent. As applications pile up, teams fall back to skimming résumés and gambling on gut feel.
How it works
Inside the evaluation
One model guessing at a score is a black box. Truly compiles a session in passes: specialized agents extract what happened, evidence is assembled, and only then does anything render judgment.
01 — Extract
No agent scores anything. They emit structured events, so a weak one can be swapped without touching the rest.
Interface
Reads what's on screen
Session
Watches the work unfold
Code
Tracks what changed
AI work
Follows the exchange
02 — Assemble
A timeline builder merges every stream into a single ordered account, then behavior detection infers what the candidate was actually doing.
Inferred behavior
Stored as permanent, queryable evidence — the source every score is later held to.
03 — Judge
Each evaluator sees only the evidence relevant to it — the debugging evaluator never reads unrelated signals — then a meta evaluator synthesizes the verdict from their scores alone.
sees: notes · file opens · first edits
sees: diffs · complexity · structure
sees: errors · terminal · edits · runs
sees: test runs · coverage · outcomes
sees: prompts · accepts · rewrites
sees: messages · reviews · rationale
Why Truly
Other platforms watch the code. Truly watches the work — technical judgment and interpersonal skill, in one simulation.
A familiar environment changes how people perform. Candidates use their own machine, editor, and shortcuts — so you see how they actually ship, not how they cope with a browser IDE.
Screen recording, codebase, and AI interactions are analyzed together, so nothing between the ticket and the submission is a black box — how they plan, debug, and decide.
Simulated workplace scenarios — a pushy deadline, a disagreeing teammate — score communication and collaboration with structure, not vibes.
Because simulations are generated from the job description and graded on process, the same approach extends to non-technical roles as your team grows.
Pricing
We scope hiring volume, roles, and integrations with each team, then quote privately.
FAQ
Truly generates a role-specific work simulation directly from your job posting — using its description, seniority level, and stack. Instead of pulling from a static library, every candidate gets a task that mirrors the actual job.
On their own machine, in their own setup. Running the assessment locally captures a far more realistic workflow than a browser sandbox — the tools, shortcuts, and environment they actually use to ship.
The full work session: prompts, actions, edits, debugging, and decision-making — plus collaboration signals like PR-style review, messaging, and handoffs. You see the whole process, not just the final output.
We measure how candidates work in modern AI-assisted environments. Every AI interaction is logged and scored for prompt quality and iteration — we grade fluency, not abstinence.
A structured, hiring-ready report with rubric-based scores and evidence, ranked across your applicant pool. Candidates keep a reusable profile of verified work artifacts they can carry across applications.
Get started
Set up in a day · Works alongside your existing pipeline