All articles

Gemini Enterprise

Which Workflows Actually Deserve an Agent

5 August 20267 min readBy Aamir Faaiz

The shortlist from the first workshop is rarely the right one. Four filters: measured frequency, whether it reaches into a system, whether errors are recoverable, and whether you can define correct at all.

The shortlist everyone starts with is the wrong shortlist

Ask any organisation which workflows they want to automate and you get a version of the same list: answering repetitive internal questions, triaging inbound requests, summarising long documents, drafting first-pass responses.

It is a reasonable list. It is also the list of things that are easy to describe, which is not the same as the list of things worth building. Two of those four are usually better solved by fixing the document that keeps generating the question, and one of them is worth more than the other three combined.

The filter we use has four questions, and a workflow needs a good answer to all of them.

How often, and by how many people

Frequency is the obvious one, and it is still the one most often estimated rather than measured.

The number that matters is not how long the task takes. It is occurrences multiplied by people. A twenty minute task that four hundred people do twice a week is a far better candidate than a two day task that one specialist does monthly, even though the second one feels more painful to whoever does it.

The specialist task is usually the one that gets proposed, because that person is in the room and can describe their pain vividly. The four hundred people are not in the room. Go and count.

Does it need to reach into a system, or just read

This is the question that separates a genuine agent from an expensive search box.

If the work is answered entirely by knowing something, you need good grounding and you may not need an agent at all. If completing the task requires touching a system, opening the ticket, updating the record, checking the entitlement, pulling the timeline, then an agent is the right shape, because the value is in the doing rather than the knowing.

This is where Gemini Enterprise genuinely changes the calculation. Because it is already connected to Workspace, Microsoft 365, Salesforce, SAP and ServiceNow, workflows that used to fail on integration cost now clear the bar. Tasks that were not worth six weeks of connector work become worth an afternoon.

Failing the fourth filter is the most common and the least noticed.

What happens when it gets it wrong

Every agent will be wrong sometimes. The question is what that costs and who notices.

We sort candidates by recoverability. If a wrong output is caught by the next person in the process anyway, the workflow tolerates an agent early. If a wrong output goes straight to a customer, a regulator or a payment, it either waits or ships with a human confirming the step.

The useful reframe here is that a workflow with a natural human checkpoint already in it is not a compromise candidate. It is the best kind of first candidate, because you get the productivity benefit immediately and you accumulate the evaluation evidence that lets you remove the checkpoint later with an argument rather than a hope.

Can you actually say what correct looks like

This is the filter that fails most often and gets noticed least, usually about five weeks in.

If you cannot write down what a good outcome is for a given input, you cannot evaluate the agent, which means you cannot tell whether a change improved it, which means you are shipping on vibes. That is survivable for a demo and not survivable for something running unattended against real records.

The practical test is to collect thirty real historical cases and have the person who does the job today mark what the right answer would have been. If that exercise is straightforward, you have an evaluable workflow. If it produces an argument about what right even means, you have found something more valuable than an automation opportunity: an unresolved process decision that has been quietly costing you for years.

The candidates that survive tend to look boring

Workflows that clear all four filters are rarely the ones that were pitched in the first workshop. They tend to be unglamorous, high volume, system-touching and already checked by somebody.

That is a feature. The demonstrable win from a boring workflow buys the credibility and the evaluation history you need for the ambitious one. Starting with the ambitious one and stalling buys neither.

Measure frequency rather than estimating it, prefer workflows that touch a system over ones that only need an answer, start where errors are recoverable, and refuse to build anything you cannot evaluate. The list that survives will be shorter and duller than the one you started with, and considerably more likely to reach production.

Gemini EnterpriseAI AgentsWorkflow AutomationEnterprise AIEvaluation

Working on something like this?

No pitch, just a practical conversation with the team that builds and operates these systems in production.

Start a conversation