Before we get into the tactical chapters, I need to give you a framework for thinking about the different levels of AI. Because the biggest mistake I see companies make isn’t choosing the wrong tool but applying the right tool at the wrong level. They’ll try to solve a systems problem with a chat prompt. Or they’ll try to build an agentic system before they’ve figured out their workflows. Both mistakes cost months.
There are three levels of AI usage that matter for go-to-market teams, each useful and each with limits. And you can’t skip levels.
That last part is the one most people get wrong.
§Level 1: Crawl with Chat (A Task)
This is where most people start. You open ChatGPT, Claude, or Gemini. You type a prompt. You get an output.
Easy-peasy.
“Write me a blog post about AI in sales enablement.” “Summarize this sales call transcript.” “Rewrite this email to sound more professional.” “Give me ten LinkedIn post ideas about B2B marketing.”
Each of these is a single task. You ask, AI answers, you take the output and do something with it. The next time you need something, you start a new conversation and ask again.
Chat-level AI is useful, and I use it every day. When I need to think through a problem, I’ll have a conversation with Claude to pressure-test my reasoning. When I need a quick first draft of something short, it saves me 20 minutes. When I need to summarize a long document, it does in two minutes what would take me thirty.
But chat-level AI doesn’t compound because you need to manually prompt each step of a process, one click at a time, every time you need that process run. For an individual contributor doing individual tasks, chat is a productivity boost. The Slack Workforce Index found that workers using AI daily are about 64% more productive than colleagues who don’t.1 That’s not nothing.
But for a growth team trying to build a go-to-market engine? Chat is incremental at best. It makes each task faster without changing the underlying architecture of how work flows through the organization.
Think of it this way: chat is like having a very fast, very smart intern who can still only handle one thing at a time. You can hand them tasks all day long and they’ll do each one well. But they need to be manually told at each stopping point to connect what they did at 9 AM to what they’re doing at 3 PM.
And it’s not the tool’s fault; it’s just doing exactly what it was designed to do. The limitation is in how we’re using it.
Most companies are here. McKinsey’s 2025 State of AI survey found that 88% of organizations use AI in at least one business function.2 But the vast majority are using it at the chat level: individual people doing individual tasks, slightly faster.
That’s “using AI.” It’s not “building with AI.” The gap between that distinction is where the opportunity lives.
§Level 2: Walk with Workflows (A Process)
This is where everything changed for me.
In early 2023, I was working full-time at Copy.ai. The company was running on GPT-2.5 and most marketers hadn’t heard of it. Copy.ai was primarily known as a free AI writing tool, the kind of thing people used to rewrite a sentence or generate a headline.
But the engineering team was building workflows, which didn’t just answer one question or complete one task. Instead, they chained multiple prompts together into a sequence, where the output of one step became the input for the next. You could define a multi-step process: take this input, do step one, pass the result to step two, transform it, pass it to step three, format the output, and deliver it.
That sounds a lot more technical than it is, but it’s more of a conceptual shift in how to think about growth than anything that requires an engineering background.
In practice, the difference is stark:
Chat level: You paste a sales call transcript into Claude and ask it to summarize the key points. You get a summary. You copy all that into a Google Doc. You manually identify the pain points or ask Chat to do it for you (with a manual prompt). You then ask it to draft a follow-up email based on those pain points. You copy that email, open your CRM, and paste it in. You go back to the transcript and look for recurring themes, all in isolation of other calls that have been had that month. Then you ask it to suggest a blog topic based on those themes.
Each step is a separate task where you serve as the connective tissue.
Workflow level: You feed a sales call transcript into a workflow. Step one extracts the prospect’s pain points. Step two maps those pain points to your pre-defined value propositions. Step three generates a personalized follow-up email. Step four creates a one-pager tailored to the account. Step five tags the recurring themes and stores them in a structured database. Step six flags any themes that have appeared across three or more calls this month and suggests a content topic.
One input. Six outputs. All of it connected and fed back into a table that any other department can pull up at any time.
In the latter model, the human’s role shifts upward:
- You designed the workflow.
- You defined the value propositions it maps against.
- You wrote the example follow-up email that it uses as a quality template.
- And you review the outputs before anything gets sent or published.
But you’re not doing the repetitive assembly work because the workflow handles all those middle-of-the-process hand offs for you.
This is the moment I started wrestling with a question that would eventually define Systems-Led Growth: what parts of going to market should be seen as a process versus what should be a task?
When you start thinking in those terms, everything shifts. You stop asking “how do I write a blog post faster?” and start asking “how do I build a pipeline where one customer conversation turns into ten assets across the full funnel, automatically?”
That’s a fundamentally different question, and it leads to fundamentally different results.
The Moment It Clicked
I remember the first time I saw a multi-step workflow produce output that no single prompt could have generated. I’d built a content pipeline that took a topic brief, researched the competitive positioning for that topic, identified the search intent, generated an outline informed by both the ICP data and the competitive analysis, drafted the article, ran it through an internal review step that checked for brand voice compliance, and formatted it for publishing.
The output wasn’t perfect, and it still needed some hefty human review. But it was structurally sound in a way that a single prompt could never produce, because each step had been informed by the step before it. The outline was better because it had competitive data, the draft was better because it had a thoughtful outline, and the formatting was better because it had review criteria built in.
That’s what multi-threaded prompting does: each step carries context forward so that the whole becomes greater than the sum of the parts.
I built the content engine at Copy.ai on this principle because it was one of the first companies that made multi-threaded prompting available for non-technical people (I remember being introduced at a conference as “the perfect simpleton to explain how workflows are accessible.” My feelings would’ve been wildly hurt if it hadn’t been so dead-on accurate). The quality of outputs I was getting from that workflow simply wasn’t possible at the chat level.
And the workflows themselves became an asset. Once I’d built a content workflow that produced consistently good output, I didn’t have to rebuild it every time. I refined it, tuned the prompts, adjusted the quality criteria, but the infrastructure was there. Every new article produced was cheaper (in time and effort) than the one before it because the system was already built.
This is where the Iron Triangle starts to bend. The system handles the production, the human handles the judgment, and the cost per unit of output drops over time because the infrastructure compounds.
Workflows Across the GTM Stack
Here’s what workflow-level AI looks like across the go-to-market functions this book covers. Each gets a full chapter in Part Two; this is the overview.
1. Content production: A keyword goes in, gets researched across Google for top ranking articles, and the top three URLs are extracted and analyzed for content gaps. The keyword is then searched across our current content to find relevant internal links, and a research agent finds high-quality external links to bolster our argument. All of this is used to create a content brief. With that brief, the workflow researches competitors, generates a formal outline, drafts the article, checks it against brand voice guidelines, annotates changes needed at the sentence level (not a full rewrite), and formats it for the CMS. Human reviews, edits, and approves.
Five articles per day, one person, and codified quality in outputs.
2. Sales outbound: Account name goes in. The workflow pulls company data, recent news, hiring signals, and tech stack information, then maps the account’s profile to your value propositions and generates a personalized outreach sequence. Human reviews for tone, accuracy, and the fine line between personalized and pushy.
3. Inbound processing: Lead submits a form. The workflow enriches the record with firmographic data, identifies which content the lead engaged with, scores their intent, and generates a personalized response that references their actual behavior. The response goes out in minutes (speed to lead, baby).
4. Podcast repurposing: Transcript goes in. The workflow produces a thought leadership article, a LinkedIn post in the speaker’s voice, a newsletter draft, YouTube show notes, a landing page, a set of quote cards, and talking points for the sales team. One 45-minute conversation becomes ten assets.
5. ABM campaigns: Target account list goes in. The workflow researches each account, matches signals to value props, generates personalized landing pages, and creates multi-channel outreach sequences. What used to require a dedicated ABM team for 20 accounts now covers 200 with deeper personalization and a fraction of the time.
6. Case study production: Customer interview transcript goes in. The workflow extracts the narrative arc, pulls out key metrics, generates a full case study draft, a short summary version, a set of quote cards, and sales talking points. Human reviews and polishes. What used to take three weeks takes three hours.
Every one of these follows the same pattern: structured input, multi-step processing, multiple connected outputs, human review. The workflow does the assembly, the human does the judgment, and the system compounds because each output feeds back into a structured library that makes the next output better.
§Level 3: Run with Agentic AI (A System That Acts)
This is where the conversation gets interesting, and where I want to be careful about the line between what’s real and what’s hype.
Agentic AI refers to systems that don’t just execute a pre-defined sequence of steps. They make decisions, take actions, observe the results of those actions, and adjust their approach. In the language that’s been forming around this: they reason, remember, use tools, and operate with delegated authority.
The numbers tell a clear story: everyone is interested, many are experimenting, few have it working. McKinsey found that 62% of organizations are experimenting with AI agents, but only 23% have begun scaling them.3 Deloitte projects that 50% of enterprises using generative AI will deploy autonomous agents by 2027, up from roughly 25% in 2025.4 And only about 11% of organizations were actively using agentic AI in production as of 2025.5
In the go-to-market context, agentic AI looks like this when it works well:
Imagine a system that monitors your target account list continuously. It notices that one of your tier-two accounts just posted a job listing for a “Director of Revenue Operations,” raised a Series B round last month, and had three employees engage with your content this week. The agent scores this as a buying signal, promotes the account from tier two to tier one, triggers a personalized outreach sequence, generates a landing page tailored to the account, alerts the AE assigned to that territory, and creates a meeting prep brief with competitive positioning. No human told it to do any of this. It simply observed, reasoned, decided, and acted.
That’s agentic AI, and it’s real. Platforms like 6sense, Salesforce’s Agentforce, and others are building toward this. Some of it is working in production right now.
But there’s a lot of breathless writing about agents that doesn’t match what I’m seeing on the ground.
§Where Agents Actually Are in 2026
As of right now, in early 2026, agentic AI in go-to-market is powerful but not autonomous.
What works well: Agents that operate within tightly defined parameters on structured tasks:
- Monitoring account signals and adjusting lead scores.
- Routing inbound leads based on enrichment data and predefined rules.
- Automating the research-to-outreach pipeline where the inputs are clean and the value prop library is well-structured.
- Summarizing and routing sales call insights.
These are real, production-ready applications that deliver measurable value.
What’s emerging but not reliable: Agents that make complex strategic decisions.
- Autonomous content publishing without human review.
- Multi-step campaign orchestration where the agent decides messaging, targeting, channels, and timing with no human checkpoint.
- Agents that interact directly with prospects in ways that require nuance, empathy, or brand judgment.
These exist in demos and pilot programs, but I haven’t seen them work consistently at scale in production, and the people I trust who are building in this space say the same thing.
Granted, this could all likely change in a matter of weeks with a new update from any one of the major AI platforms.
What’s hype (for now): Fully autonomous go-to-market engines that run without human oversight. “Set it and forget it” AI marketing. Agents that replace the need for strategic thinking. If someone tells you they’ve built this, they’re either operating in a very narrow niche or they’re exaggerating.
The failure rates back this up. Gartner projects that over 40% of agentic AI projects will fail by 2027 because legacy systems can’t support the execution demands.6 Forrester estimates that three out of four firms that try to build advanced agentic architectures independently will fail.7
Those aren’t pessimistic numbers, they’re just honest ones. And they point to the same conclusion: the infrastructure has to exist before the agents can work.
§You Can’t Skip Levels
This is the thing I keep telling people, and it’s the thing they most want to argue with me about. You can’t jump from chat to agentic AI. Well, you can, but you really shouldn’t.
The workflow layer matters because the infrastructure built at Level 2 is what makes Level 3 possible.
Agentic AI systems need three things to work: clean data, structured workflows, and clearly defined parameters for decision-making.
- An agent that monitors account signals needs those signals flowing into a structured database.
- An agent that generates personalized outreach needs a well-defined value prop library to draw from.
- An agent that routes and scores leads needs enrichment workflows that produce consistent, structured output.
All of that is Level 2 work (i.e. plumbing).
Companies that try to jump straight to agents without building the workflow infrastructure first end up with what I call autonomous chaos. The agent acts, but it acts on bad data. It generates outreach, but the personalization is wrong because there’s no structured value prop library to map against. It scores leads, but the scores are meaningless because the enrichment data is incomplete or inconsistent.
I’ve seen this happen more and more lately. A company gets excited about agentic AI, buys an agent platform, gives it access to their CRM, and turns it loose. The agent starts sending outreach to accounts that don’t match the ICP. It generates follow-up emails that reference the wrong pain points because the call transcript data isn’t structured. It promotes accounts to tier one based on signals that look good in isolation but don’t actually correlate with buying intent because nobody did the work of defining what a real buying signal looks like for their specific market.
The agent did exactly what it was designed to do: it observed, reasoned, and acted. But it did all of that on top of a broken, fuzzy foundation.
This is why the book is called Pipes Before Chocolate. The chocolate (the agentic magic, the autonomous personalization, the AI that runs your GTM while you sleep) doesn’t flow without the pipes (the workflows, the structured data, the connected systems, the human-defined quality standards).
Build Level 2 first.
- Get it working.
- Get the data clean.
- Get the workflows producing consistent output.
- Get the human review loops dialed in.
- Get the content library structured and tagged.
Then, and only then, start layering Level 3 on top.
If you’re reading this book as a skeleton-crew operator, here’s the good news: Level 2 is where the vast majority of the value lives right now. You don’t need agentic AI to reshape the Iron Triangle. Workflows connected to each other, producing compounding outputs, with a human in the loop on quality; that’s enough to make one person as effective as a department.
Agents will make it better, eventually. But the foundation comes first.
§Defined vs. Decides: How to Choose the Right Level
When I’m evaluating a go-to-market problem, I run it through a simple filter. It’s not sophisticated, but it keeps me from over-engineering things or under-building them.
The filter comes down to two words: defined and decides.
If the process is defined, build a workflow.
Can you write down the steps? Are the inputs predictable? Does the same sequence apply every time, with different data flowing through it?
That’s a workflow. Content production, podcast repurposing, post-call follow-up generation, lead enrichment, case study extraction. These are processes with a known sequence. If a human could write an SOP for it, it should be a workflow.
If the process decides, consider an agent.
Does the system need to observe something, evaluate it against context, choose what to do, and then act differently depending on what it found?
That’s agentic territory. Account signal monitoring where the system sees a hiring signal, evaluates whether it’s relevant to your ICP, promotes the account, and triggers outreach. Dynamic lead scoring that adjusts routing based on signals changing in real time. Competitive intelligence that detects a pricing change and decides which accounts need to hear about it.
One-time task with a clear input and output? Use chat.
Summarize this document. Rewrite this paragraph. Brainstorm ten ideas. In other words, don’t build a workflow for something you’ll do once.
The mistake I see most often: companies trying to make agents do workflow work. They build an “autonomous content agent” when what they need is a content production workflow with a human review gate. The agent adds complexity, unpredictability, and failure modes without adding value, because the process was known all along. They just wanted it to sound more impressive.
Agentic is the hype (and for good reason), but workflows are where the compounding happens. The companies that get ahead with AI will be the ones who had the best-defined processes, the cleanest data, and the deepest institutional knowledge baked into their infrastructure before handing it off to agents.
Most skeleton-crew operators should spend 90% of their time building Level 2 systems right now. Maybe 5% on chat for one-off tasks. And maybe 5% exploring Level 3 for the one or two use cases where their workflows are mature enough to support it.
The Part Two chapters in this book are almost entirely Level 2 playbooks, because that’s where the highest-impact work is for a small team in 2026. I’ll note where agentic capabilities are ready to layer on, but I won’t pretend that Level 3 is where most readers should start.
Start where the value is: pipe before hype.
§The Proprietary Tools Wave
There’s a shift coming that I want to flag because it changes how you should think about infrastructure.
In 2023 and 2024, individual team members started using ChatGPT accounts for one-off tasks, often without telling anyone. Marketing wrote blog drafts, sales summarized calls, and product brainstormed feature specs. Each person used AI independently, with their own prompts and their own understanding of the brand, the ICP, and the product positioning.
The next version of that pattern is already emerging. Tools like Claude Code make it possible for individual team members to build their own proprietary tools: custom scripts, internal dashboards, data processors, personalized workflow automations. Not just prompts, but actual tools that really do feel like magic.
- A sales ops person builds a deal scoring calculator.
- A content marketer builds an article outline generator.
- A customer success manager builds a QBR prep tool.
This is genuinely powerful because each of those tools compounds the individual’s productivity. But the problem is the same one that plagued the chat era, magnified: data silos. If every team member builds tools from their own understanding of the brand, the messaging, the ICP, and the product positioning, you end up with twenty tools encoding twenty different versions of the truth.
The solution connects directly to the infrastructure this book is about.
Build a proprietary knowledge layer (what I call a Brand Brain) that any tool can connect to. This is the structured content library from Chapter 12 extended to include your messaging framework, your ICP definitions, your mission, vision, and values, your product truths, and your quality standards. Everything your team needs to build on top of, stored in a structured, queryable format.
When a team member builds a custom tool, that tool pulls from the Brand Brain. It inherits the right messaging and references the right ICP. It will even natively use the correct product positioning. The tool is proprietary and purpose-built for their function, but the foundational context is shared and consistent.
Without the Brand Brain, the proprietary tools wave fragments your go-to-market. With it, every tool your team builds amplifies the same strategy.
I’ll come back to this in Chapter 12 (where the Brand Brain lives as infrastructure) and Chapter 13 (where the audits that populate it get done). For now, the point is this: the infrastructure you build at Level 2 doesn’t just power your workflows. It becomes the foundation that everything your team builds on top of, whether you planned for it or not.
§Where This Is All Going
I want to give you a brief sketch of the trajectory as I see it, not because it changes what you should build today, but because understanding where things are headed helps you build infrastructure that won’t need to be torn down in eighteen months.
Right now, in early 2026, the AI picture for go-to-market looks like this: chat is mature and ubiquitous, workflows are mature but underadopted, and agents are exciting but early.
Most companies are at Level 1 thinking they’re at Level 2 with C-suite demanding they dive headfirst into Level 3.
By late 2026 or 2027, I expect workflows to become the default way teams interact with AI for repeatable tasks. The companies that built workflow infrastructure in 2024 and 2025 are already seeing the compounding effects; the ones that didn’t are feeling the gap widen. Gartner’s prediction that 40% of enterprise applications will include task-specific agents by end of 20268 suggests that agentic capabilities will increasingly be embedded in the tools teams already use, rather than requiring custom builds.
By mid-2027, I think the line between workflows and agents will start to blur. Workflows will get smarter (adjusting their own parameters based on results) and agents will get more reliable (operating within broader boundaries with fewer errors). The human-in-the-loop won’t disappear, but the loop will get wider. Instead of reviewing every individual output, you’ll review the system’s performance weekly and adjust strategy monthly.
But one thing doesn’t change across any of these timeframes: the quality of the output is still bounded by the quality of the input. If your value prop library is weak, even the best agent will generate weak outreach. If your content strategy is undefined, even the most sophisticated workflow will produce content that doesn’t connect to your ICP. If your customer language database is empty, every piece of personalization will feel generic.
The infrastructure matters at every level. The pipes are always required.
What changes is how much the system can do on its own once the pipes are in place.
§What I’m Still Figuring Out
I want to close this chapter with something honest, because I think the AI conversation has too much certainty in it from people who are making educated guesses dressed up as psychic-level predictions.
I don’t know exactly where the boundary between Level 2 and Level 3 should be for a skeleton-crew operator. I know the principle (build workflows first, add agents later), but the specific point where you should start experimenting with agentic capabilities varies by team, by industry, by data quality, by budget, and by risk tolerance.
For my own work, I’m still primarily in Level 2, now dabbling with Level 3. The workflows I’ve built for content production, AEO optimization, and account research are where 95% of my value comes from. I’ve experimented with agentic approaches for monitoring account signals and competitive intelligence, and the results have been promising but inconsistent. Some days the agent surfaces an insight I wouldn’t have found. Other days it surfaces noise that wastes my time.
I think the honest answer is that for most operators reading this book, Level 2 is where you should live for the next 12 to 18 months. Build the workflows, at least in the experimental phase.
Then, when your workflows are producing consistent output and your data is clean and structured, start experimenting with agents in narrow, well-defined use cases where the cost of an error is low.
That’s not as exciting as “deploy an autonomous growth engine next week.” But it’s true. And I’d rather give you something that works than something that sounds impressive in a LinkedIn post.
§Before the Tactical Chapters
There’s one more foundational idea before we get into Part Two: the Human-in-the-Loop Principle.
Every workflow in this book, every system I describe, has a human somewhere in it. Not because I think AI can’t do the work, but because the specific things humans are good at (knowing what matters, knowing when to stop, knowing who to talk to, saying things only they can say) are exactly the things that make the difference between output and impact.
The next chapter defines where humans fit in the system, what they should do, what they should let the system do, and the principle that governs all of it: good in, good out.
Then we lay the pipes.
Notes
- [1] Slack, “Workforce Index,” 2024–2025. Research on daily AI users’ productivity gains. ↩
- [2] McKinsey & Company / QuantumBlack, “The State of AI: How Organizations Are Rewiring to Capture Value,” March 2025. ↩
- [3] McKinsey & Company / QuantumBlack, “The State of AI,” March 2025. Data on agent experimentation (62%) and scaling (23%). ↩
- [4] Deloitte, “Agentric AI” research, 2025. Projection on autonomous agent deployment by 2027. ↩
- [5] Deloitte, “Agentic AI” research, 2025. Data on production-level agentic AI adoption (11%). ↩
- [6] Gartner, 2025. Projection on agentic AI project failure rates by 2027. ↩
- [7] Forrester, cited in agentic AI industry research, 2025. Estimate on independent agentic architecture failure rates. ↩
- [8] Gartner, cited in multiple 2025–2026 enterprise AI analyses. ↩