What AI Agents Can and Cannot Do — A Back-Office Breakdown

A mixed pile of paper forms being sorted into four trays, with a single sheet set aside and held back

“We understand what an AI agent is. We still don’t know what it would do in our company.” That is where most evaluations stall.

This article breaks it down by business function, and then — more usefully — says plainly what is still out of reach.

The pattern that works: describable steps, inconsistent formats

An AI agent pays off when three things are true of a process:

  • You can explain the steps to a new hire
  • The formats vary, which is why conventional automation never covered it
  • There is real volume — a few hours a week or more

The corollary matters as much: if you cannot explain the steps, it cannot be automated. That has nothing to do with how good the technology is. Nobody can automate an undefined task.

By business function

Accounting and finance

TaskStatus
Reconciliation of invoices against delivery notesWorks. Differences in amount and quantity are surfaced; only the mismatches go to a person
Consolidating and totalling many PDFsWorks. Mixed layouts are absorbed
Catching double entries and tax errorsWorks, given rules defined up front
Final approval of journal entriesPerson decides. The agent proposes; a person confirms

Production control

TaskStatus
Digitising handwritten production reportsWorks. Accuracy varies with the condition of the form
Automatic totals by line and by processWorks, in the shape of your existing summary sheet
Flagging days outside thresholdWorks, once thresholds are set
Diagnosing the cause of a defectHard. The data can be assembled; the judgement is human

More on this in production report digitisation.

Sales admin and order processing

TaskStatus
Digitising faxed purchase ordersWorks. Handwriting and faint scans vary
Checking for duplicate orders and credit limitsWorks
Drafting delivery-date repliesWorks, on the assumption a person reviews before sending
Deciding on price negotiationsDo not hand this over. Accountability becomes unclear

HR and general affairs

TaskStatus
Digitising business cardsWorks
Searching and summarising internal policyWorks — but design it to always cite the source passage
Sorting inbound application formsWorks
Evaluation and hiring decisionsDo not hand this over

What is still hard

This is the part most vendors leave out, so here it is directly.

Badly degraded handwriting. If a person has to squint and guess, so does the software. What matters is not pushing accuracy higher — it is designing the system to detect that it could not read something and escalate it. An unreadable field is a minor problem. An unreadable field silently recorded as a number is a serious one, because it enters your totals and nobody sees it again.

Judgements without precedent. An agent can retrieve similar past cases and propose an answer. It cannot own a decision that has no precedent to reason from. Build on the assumption that a person makes the call.

Processes with no stated purpose. “We’ve always done it monthly” is not a specification. If nobody can say what the output is for, the work is not to automate it — it is to find out whether it should exist at all. This surfaces more often than people expect, and it is not a technology problem.

Anything where accountability sits with a person. Credit, hiring, discipline, safety checks, statutory inspection. The limit here is not capability, it is responsibility, and no improvement in the technology moves it.

One more, quieter than the rest: accuracy drifts. Forms get revised, suppliers change their layouts, a new plant sends different paperwork. A system that was right in month one is not automatically right in month twelve. If nobody owns accuracy monitoring, the rollout degrades quietly rather than failing loudly.

Finding your first process

Filter internal candidates in this order:

  1. List every process involving manual re-keying. A person copying values from one system into another — Excel counts. This is the richest source.
  2. Keep the ones taking more than two hours a week. Below that, the setup cost outweighs the return.
  3. Keep the ones where mistakes get caught downstream. An existing check step makes a first rollout safe.
  4. Start with the ones whose data can leave your network. Processes handling data that cannot leave require a closed-network setup and more decisions up front. Prove the value on something simpler first, then extend.

Summary

  • What works today: manual re-keying, reconciliation, sorting, and recurring reports — processes with describable steps and inconsistent formats
  • Design decisions carrying real accountability to stay with people
  • Still hard: badly degraded handwriting, judgements without precedent, processes with no stated purpose — and accuracy drift that nobody owns
  • To choose a first process: manual re-keying → over two hours a week → mistakes caught downstream → data can leave your network

The fastest way to know which of your processes qualifies is to look at the actual forms and the actual volumes. We run requirements definition, data handover, and verification remotely, so evaluation does not depend on where you are. A 30-minute consultation is enough to name candidate processes and estimate the hours they would save.

FAQ

Can it read handwritten forms?

Usually yes, though accuracy depends on the condition of the form. The question that matters more is whether the system detects what it could not read and routes it to a person. A confident wrong reading is more dangerous than a blank one.

Do we have to change our existing forms or report formats?

No. The normal approach is to leave both the entry format used on the floor and the summary format you report on exactly as they are, and automate the steps in between. Changing the floor to suit the software rarely sticks.

How do we choose which process to start with?

Filter in this order: manual re-keying is involved, it takes more than two hours a week, mistakes get caught downstream, and the data can leave your network. Processes that fail the last test need a closed-network setup, which adds decisions — start elsewhere and you will see results sooner.

Are there processes we should never hand over?

Yes. Anything where a person must own the final call — credit decisions, hiring, discipline — and anything where a mistake is immediately a safety incident, such as safety checks or statutory inspection. An agent can draft a recommendation; the decision stays with a person.

Do we need to be in Japan to work with you?

No. Requirements definition, data handover, and verification are all run remotely by default. On-site work is arranged separately when a project genuinely needs it.

Find out how much of your workload an AI agent could take over.

Tell us about the work on a 30-minute call. We will tell you what is achievable, what is not, and roughly what it costs.

Book a 30-minute call →
Contact →