How a small talent team should roll out an AI hiring tool in 90 days

A 90 day plan for small in-house talent teams: start on one live vacancy, capture your baseline first, split the work with the agent, then widen on evidence.


Roll it out on one live vacancy, not a pilot. Capture your baseline in week one: time to shortlist, sourced response rate, hiring manager satisfaction. Give the agent sourcing, enrichment and first drafts. Keep intake, judgement and the send button with the recruiter. Add a second similar role by day 60, then widen or stop on the numbers at day 90.

This one is for in-house teams

Three pieces on this blog already cover AI rollout, all written for people who sell recruitment: the two week guide if you are solo, the multi-client playbook if you run client accounts, the enablement guide if you are training a team across accounts.

This one assumes something different. You are one to five recruiters inside one company. Your hiring managers are colleagues you will see next week, not clients you can fire. You have one employer brand, so a bad outreach week follows you around. And you have a dozen live vacancies nobody will let you pause. That last constraint is the whole design problem: you cannot roll anything out by stopping work.

Start with a vacancy, not a pilot

The instinct is to convene a small group, agree evaluation criteria, and run a three month proof of concept. Resist it. MIT’s NANDA initiative found around 95% of enterprise generative AI pilots delivered no measurable impact on profit and loss (Fortune’s write-up of the GenAI Divide report). The failure they describe is not bad models, it is tools that never get wired into how work happens.

A pilot has no owner and no deadline. A requisition has both.

So pick a role you hire repeatedly, with enough volume to show a pattern within six weeks. Not your hardest search, not the executive backfill, not the role your CEO asks about daily. Then pick a second, similar requisition and run it the old way. That is your control. Without it, every improvement you report in month three gets argued away as “the market got easier”.

The five numbers to capture before you change anything

This is the step teams skip, and the only one you cannot redo. Once the agent is in the workflow, the workflow without it is gone. Half a day, week one.

NumberWhere to get itWhy it matters
Days from intake to first shortlistATS stage timestamps, last 10 closed reqsThe first thing an agent should move, and the one hiring managers feel. Not the same as time to fill
Response rate on sourced outreachYour outreach tool or inbox, last 200 first messagesGuards against volume replacing quality
Shortlist to first interview conversionATS, per hiring managerIf it drops, the agent is surfacing candidates you like and your hiring managers do not
Recruiter hours per requisitionOne honest week of time tracking, not an estimateTime saved is the return you will be asked to prove
Hiring manager satisfaction with the slateOne question, 1 to 5, after every shortlistThe cheapest quality signal you will ever collect

Two of the five are already in your ATS. To turn the hours into a figure finance recognises, the ROI calculator does that arithmetic.

Who does what

Agree this before day one, not by argument in week five.

The agent doesThe recruiter doesNobody delegates
Searching the market, building the long listRunning the intake and writing the brief the agent works fromThe decision to reject a candidate
Enrichment, deduplication, ATS cross-checksReading the reasons behind a ranking, and challenging themThe hard trade-off conversation with the hiring manager
Salary and availability benchmarking for the briefChoosing who goes on the shortlistAnything a candidate would want a human to have decided
Drafting outreach and follow upsEditing and sending those messagesThe final hire or no hire call

One clarification on that last row, because vendors are careless about it. Avery drafts LinkedIn messages and the recruiter sends them. If a tool says it will handle LinkedIn outreach, ask who presses send.

Three ways this goes wrong

You trust a ranking you cannot see the reasons for. Test this in week one: on a role you know cold, ask the tool why it ranked the top five candidates highly and the bottom five poorly. If the answer is a number and a progress bar, you have bought a black box, and you will over-trust it or ignore it within a month. If it names the evidence, you can argue with it, which is the point. That is explainable matching in practice, and what turns sceptical hiring managers into allies rather than blockers.

You automate outreach until it reads like spam. Pin, a recruiting outreach vendor, analysed more than 5 million AI drafted messages sent through its platform between January 2024 and June 2026: AI drafted cold emails earned a 4.97% reply rate, recruiters’ hand typed first messages 12.6% (the study). Treat it as directional: four million automated sends against five thousand hand written ones that went to better chosen people. The direction is still the warning. You have one employer brand and a candidate market that talks to itself, and volume is the easiest thing an agent gives you.

You skip the intake. An agent inherits whatever brief you hand it, so a vague requirement does not stay vague, it scales. Teams that get value in 90 days spend more time on intake after the rollout, not less, because the brief is now an input to a system rather than a note in someone’s head. If yours is a fifteen minute chat and a job description copied from last year, fix that first. It is free, and nothing else works without it.

When to widen, and when to stop

Set the gates now, while you are still calm about it.

Days 1 to 30. One live requisition, one control, baseline captured, work split agreed. Success looks like the recruiter still doing their job, with the long list arriving faster.

Days 31 to 60. Add a second requisition in the same role family. Not a different family, not a second team. You are testing repeatability, not breadth. Keep the control running.

Days 61 to 90. Decide: widen, narrow, or stop. Write the kill criteria in week one so this is not a debate about who championed the tool. No improvement in days to first shortlist, a fall in hiring manager satisfaction, or a response rate below baseline: any one of those and you fix the input or you stop.

Then widen by role family, never by headcount. Moving from engineering to commercial roles changes what the agent is asked to judge, so it gets its own small run. Giving four more recruiters a login is not a rollout, it is a licence purchase.

The compliance step you cannot skip

Recruitment AI sits in Annex III of the EU AI Act, and the timeline moved this summer. Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force on 27 July 2026 and pushed the application date for standalone high risk systems from 2 August 2026 to 2 December 2027 (Lewis Silkin’s summary).

More time, not a pass. The Article 4 AI literacy obligation already applies, and a rollout is the natural moment to meet it: record who was trained, what the tool does, where a human decides, and what candidates are told. Half a page in month one beats a policy written in 2027.

What not to break

Your ATS stays the system of record for all 90 days. Do not migrate data, do not change stages, do not let a second tool become where people look for the truth. Change one variable at a time, or you will never know which one worked.

Eurostat put AI use at 20.0% of EU enterprises with ten or more staff in 2025, up from 13.5% in 2024, but only 17% among small enterprises against 55% of large ones (Eurostat). Small teams are not late. They are the ones who cannot absorb a failed rollout, which is why the one vacancy method suits them.

Ninety days from now, the question is not whether the team likes the tool. It is whether a hiring manager got a better shortlist, sooner, and can tell you why.

If you want a second opinion on which requisition to start with, book a call, or see how this works for in-house talent teams.