Which internal processes should you automate with AI agents first, and which never?
Start with work that is high-volume, repetitive, and cheap to undo when the agent gets it wrong: ticket triage and summaries, first drafts of documentation, data entry between systems, routine reporting. Never automate anything that changes a production system, moves money, or touches a customer without a named person approving the action. The order matters more than the tool.
- Less experienced and lower-skill agents
- All agents
Experienced, high-skill agents saw "minimal impact" in the study’s words, which is why we start with the people the tool actually helps.
Source: NBER / Quarterly Journal of Economics, Generative AI at Work (Brynjolfsson, Li, Raymond) (2023). About 5,000 customer-support agents at a Fortune 500 software company, staggered rollout of a generative AI assistant.
Data table
| Item | Value |
|---|---|
| Less experienced and lower-skill agents | 35% |
| All agents | 13.8% |
What to automate with AI agents first
The four tests
Run every candidate through these before anyone opens a tool.
1. Is the error cheap to undo? A wrong ticket category costs a technician ten seconds. A wrong firewall change costs a weekend. The first automations are the ones where the mistake is caught by the next person to look and fixed in one click.
2. Is there volume? An agent that saves four minutes on a task done forty times a day is worth building. One that saves an hour on a task done monthly is a distraction until the first kind are done.
3. Can you measure it? Before and after, in the same units. Minutes per ticket, days to close the month, hours of manual reconciliation per week. If there is no baseline, there is no result, only a feeling.
4. Is there a named approver? Every action that leaves the internal boundary, changes a system of record, or reaches a customer has a person attached. Not a team. A person.
Anything that passes all four goes on the list. Anything that fails the first or the fourth does not, whatever the vendor demo showed.
What goes first, in most organizations
- Ticket triage, categorization, and summaries in the service desk
- First drafts of knowledge-base articles from resolved tickets
- Meeting notes and action items, with a person confirming the actions
- Data entry between two systems that do not talk to each other
- Routine reports that are assembled by hand every week
- First-pass responses to security questionnaires, from your own policy corpus
- Contract and policy consistency checks, flagging contradictions for review
Notice what these have in common: an agent does the reading and the drafting, and a person does the deciding. That split is not a limitation of 2026 tooling. It is the design.
What stays human, for now and probably for good
- Anything that runs with administrator rights
- Changes to firewalls, backups, identity, or production databases
- Sending money, approving invoices, changing pricing
- Communication to a customer about an incident
- Hiring, performance, and termination decisions
- Anything a regulator would ask you to explain later
An agent can prepare every one of these for a person. It should not perform them. The organizations that got this wrong in 2025 and 2026 did not fail because the model was bad; they failed because nobody was accountable for what it did.
Where the gains actually land
The largest field study so far put a generative AI assistant in front of about 5,000 customer-support agents at a Fortune 500 software company. Issues resolved per hour rose 13.8% on average, but 35% for the least experienced agents, with minimal effect on the most experienced. Attrition among agents with the tool was 8.6% lower.
Two lessons for choosing the first automation. Start where the ramp is steepest, because that is where the return is. And expect the constraint to move: once triage is automated, the bottleneck becomes senior review, and the plan has to account for it.
The MSP version
For managed service providers the list is shorter because the work is more uniform. Kaseya’s 2026 survey of more than 1,000 MSPs found 53% already using AI to automate ticketing, patching, and monitoring, while more than half have automated only about a quarter of their workload. The remaining three quarters is mostly the same tier-1 work, and the first tests above apply to it directly: triage, summaries, documentation, and runbook drafts first; anything that touches a client system only with an approver and a log.
How to sequence it
- Pull ninety days of the work: tickets, requests, reports, whatever the unit is. Categorize and time it.
- Rank by hours consumed multiplied by how cheaply an error is undone.
- Take the top three. Build the narrowest possible version of each.
- Measure for four weeks against the baseline.
- Only then widen the scope or add the next three.
Three automations that work beat twelve that were announced.
Sources
If this is your situation
The three-week assessment is where we work out which of these applies to you, in writing, before anyone builds anything. Talk to us.