
You've done the interviews. You've seen the portfolio, called the references, watched someone reverse a linked list on a whiteboard. Then you signed, and three weeks in you learned the thing no interview told you: whether they ship well on your problem, in your codebase. The interview was never going to tell you that; it measures proxies. There's one test that measures the actual thing — and it's the real way to evaluate a dev subscription: run it on a single ticket.
Key takeaways
A one-task trial is how you evaluate a dev subscription before committing: you pay the vendor to deliver one real, scoped ticket from your backlog and grade the shipped work against a rubric you set in advance, measuring output on your codebase instead of the proxies an interview measures.
Look at what an interview actually grades. A portfolio is their best past work, curated, on a problem that wasn't yours. References are hand-picked and coached. A whiteboard measures performance under artificial pressure: how someone codes with a marker while a stranger watches, which is not a situation that will ever recur on the job. Three proxies, none of which answers the only question you care about: will this person do good work on the thing you're going to hand them? The same gap opens whichever staffing model you're about to commit to, whether that's a hire, a marketplace, or a subscription.
The research on selection methods has been clear for a long time. Work-sample tests, giving someone a slice of the real job and watching them do it, are the single highest-validity predictor of job performance, higher than any interview [C1]. Compare that to the methods you're actually using. Unstructured interviews sit at .38. Even a disciplined, structured interview lands around .51, the same neighborhood as a general mental ability test [C2]. A real task beats five interviews because it measures output instead of the story someone tells about their output. SHRM has made the point plainly: interviews can be misleading, because many people can make a good impression for a short period, and a trial adds a layer of transparency an interview can't [C3].
None of this is exotic. It's the difference between reading a review of a restaurant and eating there.
A trial only works if the task is real. That means three things, and getting any of them wrong wastes everyone's time.
It has to be genuine work: a small feature or an actual bug pulled from your backlog, the kind of ticket you'd hand them in month two anyway. A toy problem tests nothing about your domain. "Build a to-do app" tells you the candidate can build a to-do app, which you already assumed and don't need. It has to be bounded, scoped to roughly one delivery cycle, small enough to review in one sitting, self-contained enough that they can finish it without a week of onboarding. And it has to be representative: if your real work is untangling a legacy service with thin tests, don't hand them a greenfield function and pretend you've learned something.
Write the definition of done before they start. Not after — before. The moment you leave "done" undefined, you give yourself room to move the target once you've seen the work, and then you're grading your gut, not their output. Decide what finished looks like while you're still neutral about who's doing it.
Same discipline, one level up. Decide what "good" means in advance, so you're checking against criteria instead of rationalizing a decision you've already made emotionally.
Here's what an experienced reviewer actually reads in a small pull request. Does it meet the spec. Is the code readable and covered by tests. Did they ask the right clarifying questions before writing a line. Did they flag an edge case or a risk you hadn't thought of. What was the communication cadence while the work was in flight, and how did they behave when the ticket turned out to be ambiguous, because a real ticket always is. A good reviewer is reading intent, correctness, edge-case handling, testability, and business context. They are not counting your indentation.
Notice how much of that list is invisible on a whiteboard. You cannot see how someone handles ambiguity in a problem that has a known answer. The task surfaces it for free.
Yes. Pay for any real trial task: unpaid take-homes repel strong engineers and quietly filter for candidates who can afford to donate the hours, so you end up testing free time instead of skill.
Here's where teams flinch, and it's the wrong instinct. If a trial task is such a good signal, why not just send it out unpaid to everyone?
Because unpaid trials repel exactly the people you want. Strong engineers routinely balk at unpaid take-homes, reading them, correctly, as a chore done for free — and a bad test proves nothing and upsets everyone [C10]. It's also exploitative: an unpaid task extracts real labor and piles stress onto people already navigating economic uncertainty [C12], and it filters for whoever can afford to donate the hours, loading the burden onto the financially constrained [C4]. You're not testing skill anymore; you're testing free time. Paying for the trial levels that power imbalance and removes that burden [C4]. A paid, low-stakes, real task aligns the incentives and tells you the truth.
This is established practice, not a fringe idea. Gumroad pays candidates as contractors to do the actual job rather than trusting interviews to predict it. Sahil Lavingia put it bluntly: "It's so very hard a priori to know if someone can do the job. So we just pay them to do it!" [C5]. PostHog, Linear, Automattic, and Auth0 all run paid work trials on real codebases [C6]. James Hawkins of PostHog says the trial "makes it obvious who to hire," and that it's "frequently surprising how someone performs relative to what we thought" [C6]. That surprise is the whole point — it's the gap between the interview impression and the shipped work, made visible before you're committed.
Now the honest part, because a trial oversold is just another sales pitch. One task is a real signal, but it's a narrow one, and pretending otherwise is how you get burned again.
A single paid ticket predicts first-pass output quality and communication well. That's a lot — it's most of what you were guessing at. But it does not tell you about long-haul reliability, on-call behavior under a 2 a.m. page, or how someone handles a genuine crisis three months in. A short trial is a small sample, taken outside normal working conditions. It can even work against you: the pressure of being evaluated can trigger a kind of evaluative threat, where a careful person avoids mistakes instead of exploring the problem, so you end up measuring stress tolerance rather than fit [C11]. So don't treat one task as a culture verdict. It isn't one.
The long-haul question is real, but a longer trial is the wrong answer to it: you'd just be paying more to extend a sample that's still artificial. The right answer is structural: a low-risk way to keep going or stop cleanly, so month two decides itself on real work under real conditions. That's a cancel-anytime arrangement, not a leap of faith.
Most of the signal is in how someone works the task, not just the artifact they hand back.
The green flags: clarifying questions before any code gets written. A small, clean PR you can actually review in one pass. A flagged edge case you'd missed — that one is gold, because it means they were thinking about your problem, not just closing a ticket. And an honest "this will take longer because X" instead of a silent overrun.
The red flags mirror them. Silence followed by a big-bang delivery you have to reverse-engineer. No questions at all, which is its own tell, because nobody understands a real ticket without at least one. Scope creep that wanders off the definition of done you wrote up front. And defensiveness when you give feedback, because you will give feedback, and how someone takes one round of review tells you more than the code ever did.
Collapse the whole thing down. One real task, paid, scoped to a delivery cycle, judged against criteria you wrote before you saw the work. That is the entire risk you're taking. Set it next to the conventional alternative — a technical role takes a median of 75 days just to first fill [C8], and in-house recruiting runs $9,000 to $25,000 per hire [C7] — and the trial isn't the risky option. It's the cheap one. (The same one-ticket test works whether you're weighing a subscription or an open marketplace like Toptal or Upwork.)
This is exactly what DevOD is built around. Intake is a 15-minute fit call and a one-task Proof of Quality: you judge a real engineer's output on a real ticket before you subscribe to anything. If it lands, you get a same-day match and your first task shipped within five business days, then a three-day task cycle with daily async updates and a task-by-task approval gate — you sign off before the next one starts. A single stream is $3,495/mo for one engineer; a dual stream is $6,795/mo for two. You own the IP, signed up front, and the engineer works inside your environment and your access controls, under NDA. It's month-to-month, cancel any time, no lock-in, which is how the long-haul question actually gets answered, and how an on-demand product team is meant to work: by low-risk continuation on real work, not by a signature you can't take back. If the matched engineer isn't right, there's a five-day replacement guarantee.
You don't need to trust the pitch. You need one ticket back, graded against a rubric you wrote first. Everything else follows from that.
How do you evaluate a dev subscription before committing? Hand the vendor one real, paid ticket from your backlog, scoped to a delivery cycle, and grade the shipped work against a rubric you wrote first. It measures output on your codebase instead of interview proxies like portfolios, references, and whiteboards.
Should you pay for a trial task? Yes. Unpaid take-homes repel strong engineers and filter for whoever can afford to work for free, so you end up testing spare time rather than skill. Paying levels the power imbalance and gets you an honest signal [C10][C12][C4].
What makes a good trial task? It has to be genuine work pulled from your backlog, bounded to about one delivery cycle and reviewable in one sitting, and representative of your real codebase — not a toy problem like "build a to-do app."
Can one task really tell you if a developer is good? It predicts first-pass output quality and communication well, which is most of what interviews only guess at. It does not tell you about long-haul reliability, on-call behavior, or how someone handles a crisis months in [C11].
How is this different from a take-home interview test? A take-home is usually an unpaid, artificial exercise; this is a paid, real ticket from your backlog judged against a rubric written in advance. The signal comes from real work under real conditions, not a contrived puzzle.