Where to Point AI First
Most AI pilots don't fail because the tech is hard. They fail because they were pointed at the wrong work. This is the selection framework I use — industry × function × company size, scored by frequency and value — from running five AI-native companies with zero hired employees.
Table of Contents
There is a stack of guides — five of mine included — on how to build an AI-native system: the context layer, the closed loop, the evals, the skill chains. Almost nobody writes about the decision that comes first and decides whether any of it works: where to point it.
I run five companies across five different industries with co-founders, AI agents, and zero hired employees. The single most important call I make in each one isn't which model or which framework. It's which workflow gets the agents first. Get that right and the system compounds. Get it wrong and you join the 95% of pilots that demo well and never ship.
This is the selection framework: the three vectors, the grid that ranks them, and the exact rubric I score a workflow with before I build a thing. It's the half of "going AI-native" that the architecture guides skip.
3
vectors: industry × function × size
95%
of AI pilots never reach production
5
companies, one selection method
0
hired employees behind it
Why Most AI Pilots Fail (and It Isn't the Tech)
A widely-cited 2025 MIT study found 95% of enterprise generative-AI pilots delivered no measurable P&L impact. The model was almost never the reason. The aim was.
When a mid-market company tells me their AI pilot stalled, I can usually guess what it was pointed at before they tell me. It was the impressive thing: the strategy copilot, the "ask our whole company anything" bot, the board-deck generator. Glamorous, rare, and soaked in human judgment — which is exactly the work agents are worst at and you get the fewest reps to improve.
Demis Hassabis put the principle better than I can: "Running 100 miles an hour in the wrong direction is worse than standing still." Speed is the whole promise of an AI-native system — produce anything in minutes — but speed in the wrong direction just gets you to the wrong place faster, and burns the organization's patience on the way. The companies that succeed don't run faster than the ones that fail. They aim better.
The production gap is real, but "it didn't make it to production" is a symptom. The disease is that the first workflow was unwinnable. Pick a function an agent can nail, on a cadence frequent enough to tune, and the same team that failed the first pilot ships the second one. This whole guide is about making that pick on purpose instead of by accident.
AI
The next teardown, without checking back
About one a week from the five companies I run with co-founders, AI agents, and zero hired employees — what I changed, and the part that broke.
One click to unsubscribe.
The Three Vectors: Industry × Function × Company Size
Every selection decision lives on three axes. Two of them you choose. One of them — for a mid-market operator — is already fixed.
Niching down isn't one decision, it's three. Industry (the vertical — clinics, agencies, logistics, e-commerce ops), function (the workflow — intake, routing, follow-up, QA, research, drafting), and company size (who you're serving). The trap is treating "we're doing AI" as a single bet. It's a point in a three-dimensional space, and most failed pilots never chose their coordinates — they defaulted to "everything, for everyone, all at once."
Here's the part that simplifies it for you: if you're a mid-market operator, the size vector is already set. You're a $2M–$50M, operations-heavy company. You're not a ten-thousand-person enterprise that needs a transformation office, and you're not a solo founder with no operation to systematize. That's the band where an AI Operating System pays back fastest, because you have real, repeatable volume but you don't yet have the headcount or the internal team to build the system yourself.
So the real work is the other two axes — industry and function. And of those two, one of them should almost always come first.
Why Function Is Your First Cut
The unit of leverage is the function, not the vertical. I run five companies in five different industries on one operating method — because intake, routing, QA, and research look the same everywhere.
My five companies could not look more different on the industry axis. Sena is an event concierge. Precis scores expert consensus across thousands of health videos. Gavel produces cited, framework-grounded advice. TrueStandard is a verification layer for high-stakes decisions. GameTape is an ambient observation layer for executive coaching. Five verticals, five buyers, five vocabularies.
And yet they run on one operating method, because underneath the vertical they're the same handful of functions: capture the inputs, route them, enrich them, draft the output, check the output, ship it, learn from what happened. The industry changes the words. The function is what an agent does, and the function ports. That's why "function first" beats "industry first": pick the function and your work transfers across verticals; pick the industry and you re-learn the same workflows in every new vocabulary.
For a mid-market operator the same logic applies inside your own company. Don't start by asking "what's our AI strategy for the business." Ask "which function is bleeding the most capacity right now" — the one your best people keep getting pulled into, the one that's the same every time but still done by hand. That function is your first cut. The industry context you already have for free; the function is the thing you build.
Intake
turning messy inbound into structured records.
Routing / triage
getting the right thing to the right place.
Follow-up
the chase nobody enjoys.
Enrichment
adding the context a decision needs.
Drafting
first versions of proposals, replies, reports.
QA
checking work against a standard.
Research
synthesizing what's already known.
The Frequency × Value Grid
Once you have candidate functions, rank them on two axes: how often they run, and how much each run is worth. The start-here quadrant is not the one most people reach for.
The instinct is to start with the highest-value workflow you can think of. It's the wrong instinct, because value-per-run and your ability to make an agent reliable point in different directions. Reliability comes from repetition — you need many runs to tune a loop, measure it, and trust it. So the first axis that matters isn't value. It's frequency. Here's how the four quadrants play out:
| High frequency | Low frequency | |
|---|---|---|
| High value | The crown jewel. Your core revenue or service workflow. Huge, but do it second — after you've proven the loop on something lower-stakes. Failing here is expensive. | The big one-off. Board decks, M&A analysis, annual planning. High ROI per instance, but you can't tune a loop on four reps a year. Assist here; don't automate yet. |
| Low / mid value | Start here. The boring daily work — intake, triage, routing, status updates, first drafts, QA checks. Low judgment, high process, runs constantly. The reps make the agent reliable and the win compounds. | Ignore. Rare and cheap. Automating it is a hobby, not leverage. |
Start in the bottom-left: high-frequency, low-to-mid value, low-judgment work. It's unsexy, which is exactly why it's available — your competitors are off trying to build the strategy bot. You bank a reliable, measured, self-improving loop on work that runs a hundred times a week, the organization watches it work, and trust compounds. Then you move up to the crown-jewel workflow with a team that now believes, and out to the high-value, low-frequency work as assistance rather than full automation.
The order is the strategy: high-frequency to earn reliability and trust, then high-value to cash it in. Skip the first step and the second one fails in public.
How I Score a Workflow Before I Build It
The grid tells you the quadrant. This turns it into a number. Score = frequency × value × AI-fit. The highest score is almost never the workflow people first pitch as "the AI project."
I don't pick workflows on gut. I score them, and I use the same shape of formula I use to decide what to build across all five companies. When I rank market demand, the score is frequency × like-weight × operator-fit — how often a pain shows up, how strongly people feel it, and how well it fits what I'm uniquely able to do. Point that same lens inward at your own functions and it becomes:
score = frequency × value × AI-fit
- Frequency — how many times a week does it run? More reps means more chances to tune the loop until it's trustworthy. Frequency is what lets an agent compound instead of staying a demo.
- Value — capacity recovered or cost avoided per run, multiplied by frequency. A two-minute saving on something that runs two hundred times a week beats an hour saved on something quarterly.
- AI-fit — three sub-scores: judgment (lower is better), process structure (higher is better), and input legibility (can an agent read the inputs?). Low-judgment, high-process, legible work is where agents win first.
Multiply the three for a single number per function, then rank. A worked row: inbound lead intake runs ~150×/week (high frequency), recovers ~10 minutes of a coordinator's time each run (real value at that volume), and is low-judgment, high-process, and legible (high AI-fit). It scores far above "an AI strategy assistant for the leadership team" — rare, judgment-soaked, illegible — even though the second one sounds like the bigger AI project.
What gives me confidence this rubric is right isn't theory — it's that the market screams the same answer. I scored roughly 26,000 comments across the channels where this audience talks. The number-one pain, by a wide margin, was the build-to-production gap: "the building problem is basically solved; what's still broken is the execution side." The single most-liked comment in the entire scan — 508 likes — was a complaint that nobody had addressed security and data governance in a handoff. Both are high-frequency, last-mile, low-glamour functions. The demand data points exactly where the scoring rubric does: at the boring, repeatable, governable work, not the strategy bot.
The one-line version:
Rank your functions by frequency × value × AI-fit. Build the top of the list, not the workflow with the most impressive name. The impressive ones almost always score low on the two things that make an agent reliable: frequency and fit.
You Already Have the Context
A brand-new AI lab has to borrow its context. A mid-market company already owns it. That's not a disadvantage to overcome — it's your edge in choosing what to build.
A common objection to all of this is: "Sure, but you have a team of brilliant people and years of accumulated context. I'm starting from scratch." For a brand-new AI studio that's a real problem — and the answer is to borrow: pull design systems from a library, plug into a public component set, slowly load context over time. Bootstrapping from zero is a genuine constraint when you have no proprietary data of your own.
A mid-market operator has the opposite problem, which is a good problem. You're not starting from zero. You have years of operational data, a team that knows the edge cases cold, pricing exceptions and "we don't do it that way because…" rules that exist nowhere a competitor could copy. The job was never to invent context. It's to make the context you already own legible to agents — get it out of your senior people's heads and into a layer the system can read.
This is why selection is easier for you than for a startup: the function bleeding the most capacity is usually obvious to the people doing it, and you already hold the context that function needs. You're not guessing at a market. You're pointing a system at work you already understand better than anyone. That ownership is the unfair advantage — most companies just never make it queryable.
What "Good" Looks Like Once You've Picked
Selection is this guide's job. Once you've picked the function, "good" means building it as a closed loop — and I've already written that part down in detail.
Picking the right function is most of the battle, but it's worth knowing what you're aiming the pick at. "Good" is not a one-off automation. It's a closed loop: capture the inputs the function runs on, curate them so the agent sees the right context, store them in a brain the agents can search and write back to, execute the work, ship it so a customer or a team experiences it, and feed the signal from that back into the brain so the next run is better. An open loop loses information; a closed one compounds it.
I'm not going to re-teach the build here — I've written it out across the operator guides, and a reader who's chosen their function should go read the relevant one:
- What It Actually Means to Be AI-Native — the three layers (people, agents, context) and the two workflows that turn the system into speed-to-signal. Read this one first if the loop above felt abstract.
- Service-as-a-Software — the business-model lens: SaaS vs services vs service-as-a-software, the four layers, and what changes when you stop pricing human time.
- How I Run 5 Companies with Zero Hired Employees — the build order and the five operating primitives, with each company torn down to the one role it replaced.
- The Locked-Context Audit — score your own company, then see why most start at the wrong layer.
- Every Agentic Engineering Hack I Know — the ten patterns that make the loop reliable in production: context engineering, evals, subagents, self-correcting loops.
- Harness Engineering — why the scaffolding around the model, not the model, is the whole game.
- The Last Mile Is the Offer — once you've picked the function, this is how you get it into production and keep it alive: the orchestration layer, the handoff, and why the last mile is where pilots die.
- The Agent Operator: The Only AI Job That Compounds — who runs the function once you've pointed AI at it: the five-job Operator Stack, and why operating compounds while doing gets automated.
- AI Engineer World's Fair 2026: An Operator's Field Notes — 32 talks distilled into the Reliability Stack: the memory, retrieval, eval, and guardrail layers that make a pointed-at function survive production.
- The 7-Layer Reliable Agent — once you've picked the function, the seven-layer build (goal contract, verifiers, control loop, orchestration, observability, memory) that stops the agent failing silently in production.
The split is deliberate: those guides own the how. This one owns the where. You need both, in that order — aim first, then build.
Where to Start in Your Company
Name the one high-frequency function that's bleeding capacity. Score it. That's not a guess about the future — it's the first thing a real audit settles, before a line of code gets written.
Here's the exercise, end to end. List the functions your team runs every week. Score each one on frequency × value × AI-fit. Circle the highest score that sits in the high-frequency, low-judgment quadrant. That circled function is where you point AI first — and I'd bet it's not the one anyone would have written on a slide titled "Our AI Strategy."
If three different people on your team would circle three different functions, that disagreement is the finding. It means the selection isn't obvious from the inside, which is the most common reason pilots get pointed at the wrong work — everyone optimizes for the workflow they personally feel, not the one the numbers favor. Which layer and which function your company should start on isn't a vibe. It's the first thing a real audit settles, before anyone builds.
That's exactly what I do in an audit: map where the company is overstaffing work a stack of agents could do, score the candidate functions against this rubric, and hand back the one to build first and the order to build the rest. Architecture you can inspect, tied to a number — capacity recovered, cost avoided, throughput gained — not slideware.
Frequently Asked Questions
Why do most AI pilots fail?
Not because the model is weak. A widely-cited 2025 MIT study found 95% of enterprise generative-AI pilots delivered no measurable P&L impact. The pattern underneath is aim: pilots get pointed at glamorous, low-frequency, judgment-heavy work instead of the boring high-frequency workflow an agent does well. Direction beats speed.
How do I choose which workflow to automate with AI first?
Score each candidate function on three things and multiply them: frequency (how many times a week it runs), value (capacity recovered or cost avoided per run, times frequency), and AI-fit (low judgment, high process, legible inputs). Start with the highest-frequency, high-fit workflow, not the most impressive one. High frequency gives you the reps to tune a reliable loop.
Should I pick an industry or a function first?
Function. The unit of leverage is the workflow function — intake, routing, follow-up, QA, research, drafting — not the vertical. I run five companies across five different industries on one operating method, because the function ports across all of them. Pick the function before the vertical.
What if I don't have a big data set or AI team to start from?
A mid-market company already owns the proprietary context — years of operational data and senior judgment. A brand-new AI lab has to borrow context; you have to make the context you already own legible to agents. That's an advantage, not a gap, and it's what makes function selection easier: the workflow bleeding the most capacity is usually obvious to the people running it.
How do I get help picking and building the first workflow?
That's what the audit settles. A fixed-deliverable audit maps where your company is overstaffing work a stack of agents could do, scores the candidate functions, and names which one to build first and in what order. Apply at agrahri.com.
Want help aiming it?
Selection is the first thing an audit settles. I take a few a month: a fixed deliverable that scores where your company is overstaffing work a stack of agents could do, names the function to build first, and lays out the order to build your AI Operating System.
Apply for an audit