Work simulations for hiring

Hire for the work, not the interview.

Truly turns your job description into a job-specific work simulation. Candidates do the actual work on their own machine, with their own tools — and AI evaluators turn the full session into evidence you can hire on.

100+ candidates assessed · 3+ startups hiring with Truly today

  • Planninghow they broke the problem down
  • Architecturethe shape of what they built
  • Debuggingwhat they did when it broke
  • Testingwhat they checked before shipping
  • AI collaborationhow they used the tools
  • Communicationhow they wrote it up
session replay — senior-frontend · candidate 041RECORDING

Evaluation

Design process92
Code quality88
AI fluency90
Collaboration85

Strong hire signal. Plans before coding, uses AI to reason — not to outsource — and communicates risk early.

The problem

Hiring measures the performance, not the person.

AI changed how work gets done. Screening didn't keep up — so companies still gamble on candidates they've never actually seen work.

PUZZLES

Tests that aren’t the job

Generic, sandboxed challenges pose problems a candidate will never see again at work. Passing one predicts trivia recall, not on-the-job performance.

REHEARSAL

Interviews that reward polish

Behavioral rounds measure preparation, not character. Culture fit — the thing early teams live or die by — never actually gets tested.

SCALE

Grading that doesn’t scale

Human review is slow and inconsistent. As applications pile up, teams fall back to skimming résumés and gambling on gut feel.

How it works

From job description to hiring decision.

Inside the evaluation

Agents that report facts, not opinions.

One model guessing at a score is a black box. Truly compiles a session in passes: specialized agents extract what happened, evidence is assembled, and only then does anything render judgment.

01 — Extract

Each agent has exactly one job

No agent scores anything. They emit structured events, so a weak one can be swapped without touching the rest.

Interface

Reads what's on screen

Session

Watches the work unfold

Code

Tracks what changed

AI work

Follows the exchange

02 — Assemble

Facts become one coherent story

A timeline builder merges every stream into a single ordered account, then behavior detection infers what the candidate was actually doing.

12:05edited Login.tsxbuilding
12:06ran npm testtesting
12:064 assertions faileddebugging
12:14fixed, 18 passingtesting

Inferred behavior

PlanningBuildingDebuggingTestingReviewing

Stored as permanent, queryable evidence — the source every score is later held to.

03 — Judge

One evaluator per competency

Each evaluator sees only the evidence relevant to it — the debugging evaluator never reads unrelated signals — then a meta evaluator synthesizes the verdict from their scores alone.

Planning4/5

sees: notes · file opens · first edits

Architecture5/5

sees: diffs · complexity · structure

Debugging4/5

sees: errors · terminal · edits · runs

Testing3/5

sees: test runs · coverage · outcomes

AI collaboration4/5

sees: prompts · accepts · rewrites

Communication4/5

sees: messages · reviews · rationale

Meta evaluator

No score without a source.

The report ranks candidates on competency scores — and every one of them unfolds into the events that produced it. Recruiters get an audit trail, not a summary they have to trust.

Testing3/5
12:06ran npm test — 4 assertions failed
12:11read the stack trace, opened Login.test.tsx
12:14fixed the assertion — 18 passing

3 linked events · full replay available

Why Truly

Every signal, truly captured.

Other platforms watch the code. Truly watches the work — technical judgment and interpersonal skill, in one simulation.

ENVIRONMENTTheir setup, not a sandbox

A familiar environment changes how people perform. Candidates use their own machine, editor, and shortcuts — so you see how they actually ship, not how they cope with a browser IDE.

PROCESSThe thinking, not just the diff

Screen recording, codebase, and AI interactions are analyzed together, so nothing between the ticket and the submission is a black box — how they plan, debug, and decide.

CULTURECulture fit, measured

Simulated workplace scenarios — a pushy deadline, a disagreeing teammate — score communication and collaboration with structure, not vibes.

ROLESBeyond engineering

Because simulations are generated from the job description and graded on process, the same approach extends to non-technical roles as your team grows.

Pricing

Pricing is set per engagement.

We scope hiring volume, roles, and integrations with each team, then quote privately.

FAQ

Common questions

How is each assessment created?

Truly generates a role-specific work simulation directly from your job posting — using its description, seniority level, and stack. Instead of pulling from a static library, every candidate gets a task that mirrors the actual job.

Where do candidates work?

On their own machine, in their own setup. Running the assessment locally captures a far more realistic workflow than a browser sandbox — the tools, shortcuts, and environment they actually use to ship.

What exactly gets captured?

The full work session: prompts, actions, edits, debugging, and decision-making — plus collaboration signals like PR-style review, messaging, and handoffs. You see the whole process, not just the final output.

How is AI usage scored?

We measure how candidates work in modern AI-assisted environments. Every AI interaction is logged and scored for prompt quality and iteration — we grade fluency, not abstinence.

What do I get at the end?

A structured, hiring-ready report with rubric-based scores and evidence, ranked across your applicant pool. Candidates keep a reusable profile of verified work artifacts they can carry across applications.

Get started

See how your candidates actually work.

Set up in a day · Works alongside your existing pipeline