Saltar al contenido

    Audio study guide · Personal learning plan

    Go-to-market
    engineer.

    The free route of my plan, compressed into six hours a day: 9 steps over 33 working days instead of six months. Eleven tracks covering how to read the plan, the nine steps and a revision round. Each one carries a section index, the free materials and the full text of what you hear.

    11 tracks. 30 min 13 s.British English narration.
    21 September to 4 November 2026.

    Where this comes from

    The written plan

    This guide narrates the zero-cost route. The original plan holds the diagnosis, the paid-cohort version and the week-by-week calendar.

    Read the full free route ↗

    Also on YouTube

    The whole series, as a playlist

    The same eleven tracks as video, with chapters, in order, for listening on a television or in the background.

    Open the YouTube playlist ↗

    The rule that holds it together

    The first hour of every day

    Reading transcripts from my own agents, before anything else. In the six-month version it was one hour every Friday; at six hours a day it becomes daily. It is what replaces the $4,200 cohort.

    Go to the evaluations step ↗

    The accelerated calendar

    185 hours in 33 working days

    Six hours a day, Monday to Friday, from 21 September to 4 November 2026. The content does not shrink; the calendar does. A lost day is not a lost hour, it is a lost six.

    DaysDatesStepHoursDeliverable
    1 to 221 to 22 Sep1. Map and vocabulary12 hOne page per system with its approval gate marked, and the laboratory set up (DuckDB, Langfuse and the learning folder).
    3 to 1023 Sep to 2 Oct2. Evaluations45 hAn evaluation suite for Valeria and for the GoHighLevel guardian, with your taxonomy of failures. And the daily transcript hour, already a habit.
    11 to 145 to 8 Oct3. SQL and scoreboard20 hThe scoreboard with five queries you wrote yourself, running over your GoHighLevel exports, your Retell logs and the contacted file.
    15 to 189 to 14 Oct4. Observability20 hTwo Modal applications with traces, and one weekly line per agent published in Slack.
    19 to 2315 to 21 Oct5. Reading Python25 hCS50P weeks 0 to 7 finished, and three of your own scripts read, explained and each with a test.
    24 to 2722 to 27 Oct6. API, MCP and SDK22 hGhl-flow rebuilt as a governed MCP server, with free reads, a single write that leaves a proposal for human approval, and every call logged.
    28 to 3028 to 30 Oct7. Commercial systems18 hRevenue data dictionary, Score Cazador by script, deliverability runbook, and one complete circuit (list, agent, gate, appointment) measured on the scoreboard.
    31 to 332 to 4 Nov8. Experimentation15 hOne pre-registered A/B test launched, a propensity model version one compared against the Score Cazador, and the retrospective of the 33 days.
    24 to 3322 Oct to 4 Nov9. Repository and CI8 hTests running on every proposed change, protected main branch and the evaluation regression automated.

    Continuous listening

    All 11 tracks, back to back

    For the car or the walk: the whole guide in a single play, with one jump per track. The individual tracks are below.

    Listen to everything

    30 min 28 sDownload the complete guide

    Before you start · Listen to this first

    How to use this guide

    Listen

    3 min 56 sDownload to listen offline

    What the free route is in its accelerated form, the four questions that close every step, why the order is not that of a conventional syllabus, and the one rule that is not up for negotiation.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. 185 hours at six hours a day: 33 working days, from 21 September to 4 November 2026.
    2. Every step closes by answering four questions: what you learn, why it matters, what you will be able to do, and what the deliverable is.
    3. The order starts with vocabulary and evaluations, not with programming: the gap is not writing code, it is knowing whether what you built works.
    4. The route replaces a $4,200 cohort with free materials plus one discipline.
    5. At this pace the transcript hour stops being weekly and becomes daily.

    Questions to revise

    1. What are the four questions that close every step?
    2. What three things did the paid cohort give, and how does the free route substitute them?
    3. Why does the programming step come fifth rather than first?
    4. If a day collapses, what is the last thing to be sacrificed?
    Read the text you hear

    What this is and what it is for

    Study guide. Track zero. How to use this guide.

    This is the free route to becoming a go-to-market engineer, turned into material you can listen to. The plan itself tells you what to do and when. This guide is a different thing: it is for revising without a screen in front of you, while you drive, walk or wait.

    There are eleven tracks. This one, nine tracks of one step each, and a final one for revision, with questions and answers.

    Before we get to the first step, four things worth being clear about.

    The accelerated calendar and the four questions

    First. This is the accelerated version. The original plan spread one hundred and eighty-five hours across twenty-four weeks, at seven hours a week. This one packs the same hours into six hours a day: thirty-three working days, from Monday the twenty-first of September, two thousand and twenty-six, to Wednesday the fourth of November. Six and a half weeks instead of six months.

    The content does not shrink. The calendar does. That has one consequence you should hear now rather than discover later: at this pace, nothing can slip. A lost day is not a lost hour, it is a lost six.

    Second. Every step is studied with the same four questions, always in the same order. What you learn. Why it matters in your particular case. What you will be able to do that you cannot do today. And what the deliverable is, the thing that proves the step is closed. If at the end of a step you cannot answer all four, the step is not closed, however many videos you have watched.

    Why this order and not a syllabus

    Third. The nine steps are not in the order of a conventional syllabus. A syllabus would begin with programming. This route begins with vocabulary and then goes straight to evaluations, because your largest gap is not writing code: it is knowing whether what you have already built actually works.

    Code arrives at step five, and only to read and debug it, never to write it from scratch. That decision is deliberate. You already have sixty scripts running. What you do not have is a way of proving they do what they claim.

    What replaces the paid course

    Fourth, and the most important. The free route replaces a paid cohort costing four thousand two hundred dollars. That cohort gave three things the free materials do not give on their own: a calendar with dates, feedback from an instructor on your own work, and peers to compare yourself against.

    The route substitutes them like this. The calendar is the one in this plan, and it is kept. The feedback comes from Claude, but only on what you wrote first: your labels, your rubric, your judge, your query. Never written for you. And the comparison comes from your own scoreboard, which from day fourteen onwards tells you whether the agent improved or not.

    One piece cannot be substituted by anything: the hour spent reading transcripts from your own agents. In the six-month version it was one hour every Friday. At six hours a day, it becomes the first hour of every single day. That is the only non-negotiable rule in the route. Without it, the free route stops being equivalent to the paid course and turns into a list of links.

    The real cost

    A note on money. The total cost of the route is zero dollars. The only small print: the graded laboratories at DeepLearning dot AI require a subscription. The videos, which is what the route actually uses, are free. Anthropic's certifications, at one hundred and twenty-five dollars, remain restricted to the partner network: apply, and do not count on them.

    Right. Step one.

    Step 1 · Days 1 and 2 · 12 hours

    The map and the vocabulary

    Listen

    3 min 11 sDownload to listen offline

    Giving what you already do the names the industry uses: workflow against agent, the five design patterns, a well-described tool, context and the harness.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. A workflow follows steps you fixed in advance; an agent chooses its own next step, in a loop.
    2. Five design patterns: chaining, routing, parallelisation, orchestrator with workers, and evaluator with optimiser.
    3. A tool description is a prompt, not documentation: an agent that misuses a tool is almost always reading an ambiguous description.
    4. Context is a budget, not a bag: the question is not what fits but what deserves to be there.
    5. The harness is the scaffolding around the model: loop, retries, limits, logging and approval gates.

    Key concepts

    Workflow
    A fixed path. You decided the steps and the model carries them out. Cheaper, more predictable and easier to evaluate than an agent.
    Agent
    A loop in which the model picks the next action using tools, until it finishes or runs out of budget.
    Context engineering
    Deciding what goes into the model's window and what stays out, treating it as a limited budget.
    Harness
    Everything around the model that lets it survive long runs: retries, caps, logging, memory and approval gates.
    Skill
    A folder of instructions and resources Claude loads when the task calls for it, instead of keeping everything in the prompt.

    Questions to revise

    1. Of your six systems, how many are genuinely agents and how many are workflows you call agents?
    2. Which of the five design patterns do you already use without naming it?
    3. Take one ghl-flow tool: would its description tell a stranger when not to use it?
    4. Where is the approval gate in each of your systems, and who opens it?
    Read the text you hear

    In one sentence

    Step one. The map and the vocabulary. Days one and two, twelve hours.

    In one sentence: giving what you already do the names the industry uses for it.

    Workflow against agent

    The first distinction, and the most useful, is between a workflow and an agent. A workflow is a fixed path: you decided the steps in advance and the model carries them out. An agent chooses its own next step, in a loop, until it finishes or runs out of budget.

    Almost everything sold on the market as an agent is in fact a workflow. That is not a flaw: workflows are cheaper, more predictable and far easier to evaluate. Knowing which of the two you are building changes what it will cost you to run it.

    The five design patterns

    Then come the design patterns. Chaining, when the output of one model feeds the next. Routing, when a classifier sends each case down the path it belongs to. Parallelisation, when several independent runs are brought together at the end.

    Orchestrator with workers, when a central model hands out subtasks and assembles the result. And evaluator with optimiser, when a second model criticises the first one's output and forces it to improve.

    You already use four of those five without calling them by their names.

    Tools, context and the harness

    Then three ideas you will use every day. The first: a tool description is a prompt, not documentation. An agent that misuses a tool is almost always reading an ambiguous description, and the fix is not to scold the model but to rewrite the description.

    The second: context engineering. Context is a budget, not a bag. The question is not what fits, but what deserves to be there, and what should be loaded only when it is needed.

    The third: the harness. It is the scaffolding around the model, that is, the loop, the retries, the spending limits, the record of what happened and the approval gates. When a long-running agent falls over, it is almost never the model's fault: it is that the harness did not exist.

    Why it matters and what you gain

    Why it matters. Today you design by instinct and explain by analogy. That works with clients, but it collapses at two tables: in front of an engineer, and in front of a hospital's chief technology officer. With the vocabulary, you design on purpose and can justify every decision.

    What you will be able to do that you cannot today. Look at any of your six systems and say, on a single page, which circuit runs, which tools it has, where the approval gate sits and what is measured. And spot, just by reading a tool description, why an agent is misusing it.

    The deliverable for these two days: one page per system with its approval gate marked, and the laboratory set up. A question to sit with while you listen: of your six systems, how many are genuinely agents, and how many are workflows you call agents?

    Step 2 · Days 3 to 10 · 45 hours

    Evaluations and transcript analysis

    Listen

    3 min 15 sDownload to listen offline

    The longest step and the largest gap. Reading transcripts one at a time, naming the failure modes, writing binary judges and measuring whether the agent improved or got worse.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. You begin by reading real transcripts and noting in one sentence what went wrong, not by designing metrics.
    2. Those notes are grouped into named failure modes: that taxonomy is the product of the step.
    3. Each failure is measured by code where possible and by a judge, another model, where not.
    4. A judge is only useful if it agrees with your own labels: first you measure the judge, then the agent.
    5. At this pace the transcript hour is the first hour of every day, not Friday's.

    Key concepts

    Failure mode
    A concrete, repeated way the agent gets it wrong, with a name of its own. For example: books an appointment without checking availability.
    Judge
    A model that grades another model's output against a written criterion. Binary, one per failure, never a score out of ten.
    Agreement
    How often the judge matches the labels you applied by hand. Without it, the judge is an opinion dressed as a number.
    Seed scenario
    A hand-built test case representing a situation that must always work.
    Evaluation suite
    The set that runs on every change and produces a trend, not a snapshot.

    Questions to revise

    1. Why read transcripts before defining any metric?
    2. What gets measured by code, and what is left to a judge?
    3. Why a binary judge per failure rather than one overall score out of ten?
    4. How do you know your judge is any good, before using it to judge the agent?
    5. If Valeria's prompt changes on Monday, which number tells you on Friday whether it improved?
    Read the text you hear

    In one sentence

    Step two. Evaluations and transcript analysis. Days three to ten, forty-five hours. It is the longest step in the route, and rightly so: it takes eight of the thirty-three days.

    In one sentence: learning to know, with a number, whether one of your agents has improved or got worse.

    The method, in five moves

    The method has five moves, and the first one surprises everybody: you do not start by defining metrics. You start by reading transcripts, one at a time, and writing down in a single sentence what went wrong. Nothing more.

    Second move: group those notes into named failure modes. There is no universal catalogue of failures; the catalogue for your agents comes out of your transcripts. That taxonomy is the real product of this step, more than any tool.

    Third: decide what is measured by code and what by a judge. If the failure is verifiable, for instance that a phone number has eight digits or that an appointment falls inside consulting hours, that is code, and it is cheap. If the failure is a matter of judgement, for instance that the agent sounded cold or promised something it cannot deliver, that is a judge.

    Fourth: write binary judges, one per failure. Binary means the answer is yes or no, not a score out of ten, because a score out of ten cannot be audited. And before you use a judge, you measure the judge: how often it agrees with the labels you applied by hand. A judge that does not agree with you is not a metric, it is noise with decimal places.

    Fifth: build seed scenarios, run everything on every change, and read the trend.

    Why this step carries the most weight

    Why it matters. This is the largest gap against the job advert, and the one a hospital or an investor will demand before paying you anything.

    The good news is twofold. You already have the instinct: the line you wrote into the Cofepris agent, that a valid JSON does not constitute an approval, is precisely the thesis of this step. And you have the raw material almost nobody has: Retell calls and GoHighLevel conversations, real ones, in Spanish, from a specific domain. That cannot be bought.

    What you gain and how it holds

    What you will be able to do that you cannot today. Say, with a number, whether Valeria or the GoHighLevel guardian improved or got worse after a change of prompt. Hand a client, alongside the agent, the taxonomy of failures and the suite that watches it. And stop fixing agents by symptom: you will know the exact line of the transcript where it went off.

    And here is the rule that holds the whole free route together. In the six-month version it was one hour every Friday reading transcripts. Compressed to six hours a day, it becomes the first hour of every day, before anything else. That habit is what the four-thousand-two-hundred-dollar cohort would have imposed on you by calendar. Without it, this step becomes four courses watched and zero failures named.

    Step 3 · Days 11 to 14 · 20 hours

    SQL and the scoreboard

    Listen

    2 min 24 sDownload to listen offline

    Your data sits in six places. Without SQL there is no scoreboard; without a scoreboard there is no observability; and without observability the evaluations never connect to the money.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. SELECT, filters, joins, grouping, window functions and dates. That is the whole syllabus.
    2. DuckDB queries your CSV and JSON files on the Mac, with no server and no moving the data.
    3. The star schema organises the scoreboard: one fact table and several dimension tables.
    4. Flow's fact table is the funnel: message, reply, appointment, proposal, sale.
    5. The exit criterion is that you write the query, not that you ask Claude for it.

    Key concepts

    Fact table
    The events being counted, with their date and keys. At Flow: message, reply, appointment, proposal, sale.
    Dimension table
    The context you slice a fact by: contact, campaign, agent, speciality, city.
    Star schema
    One fact table at the centre with dimensions around it. It turns any business question into a short join.
    Window function
    A calculation over a group of rows without collapsing them: running totals, rankings, each doctor's first contact.
    DuckDB
    A single-file database that reads CSV, JSON and Parquet straight off disk. The laboratory for this step.

    Questions to revise

    1. What is Flow's fact table, and what are its dimensions?
    2. Which sequence produces the most positive replies, and with what query would you know?
    3. What does an appointment cost by channel?
    4. How many outreach messages have gone out today, and who sent them?
    5. Does your half past seven summary come from a query or from a file read by hand?
    Read the text you hear

    In one sentence

    Step three. SQL and the scoreboard. Days eleven to fourteen, twenty hours.

    In one sentence: to stop eyeballing your numbers and start asking them questions.

    The whole syllabus

    What you learn. The syllabus fits on one line: SELECT, filters, joins between tables, grouping, window functions and dates. That is all of it. You do not need more for ninety per cent of the questions a business asks.

    Then two practical things. How to point DuckDB at your GoHighLevel exports, your Retell logs and your file of contacted doctors, and ask them questions without setting up a server or moving the data anywhere.

    And how all of that is arranged into a star schema: one fact table, which at Flow is the funnel, that is, message, reply, appointment, proposal, sale, with dimension tables around it, that is, contact, campaign, agent, speciality and city. In that shape, any business question becomes a short join rather than a fresh script.

    Why four consecutive days

    Why it matters. Your data sits in six different places, and today every business question is either a new script or a look with the naked eye. Without SQL there is no scoreboard. Without a scoreboard there is no observability. And without observability, the evaluations from step two are left hanging: you will know the agent improved, but not that it produced one more appointment.

    In the six-month version this step ran in parallel with evaluations, three hours a week, so as not to burn out. At six hours a day that trick is no longer needed, and it would in fact get in the way: four consecutive days of nothing but queries fix the syntax far better than twenty scattered afternoons.

    What you gain

    What you will be able to do that you cannot today. Answer which sequence produces the most positive replies, what an appointment costs by channel, or how many outreach messages you have sent today and who sent them. In one line you wrote yourself, in under ten minutes, without asking Claude for the SQL.

    And have the half past seven summary in Slack come out of a query rather than a file read by hand. That is the exit criterion for this step: you write the query. Claude reviews it.

    Step 4 · Days 15 to 18 · 20 hours

    Observability and the thread to the money

    Listen

    2 min 09 sDownload to listen offline

    Today you know your agents run. You do not know what each appointment costs or where they get stuck. This step turns every run into a tree you can read.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. A trace is an agent's complete run; a span is each call to the model or to a tool.
    2. Every trace is tagged with the contact and the campaign: that is how it joins the step 3 scoreboard.
    3. A failed trace becomes an evaluation case: step 4 feeds step 2.
    4. Langfuse gives 50,000 observations a month with no card, enough for all your agents.
    5. The goal is pricing per outcome while knowing your real cost per outcome.

    Key concepts

    Trace
    The complete record of one agent run, seen as a tree of calls.
    Span
    Each branch of that tree: a call to the model or a tool, with input, output, tokens, cost and latency.
    Tag
    The business field hung on the trace (contact, campaign, client) so it can be joined to the scoreboard.
    Cost per outcome
    What it costs to produce an appointment or a positive reply, counting tokens, calls and tools.

    Questions to revise

    1. What is the difference between a trace and a span?
    2. What does a trace need before it can join the scoreboard?
    3. Last week, what did each agent do, what did it cost and what did it produce?
    4. What does an appointment produced by Valeria cost you today?
    5. How long does it take you to find the exact trace of a production failure?
    Read the text you hear

    In one sentence

    Step four. Observability and the thread that leads to the money. Days fifteen to eighteen, twenty hours.

    In one sentence: being able to say what each appointment produced by one of your agents actually costs.

    Traces and spans

    What you learn. First, two words. A trace is the complete run of an agent, seen as a tree. A span is each branch of that tree: a call to the model, a call to a tool, with its input, its output, its tokens, its cost and its latency.

    Second, how to tag. A trace without a tag is an orphaned data point. With the contact and the campaign hung on it, that same trace can be joined to the fact table from step three, and that is where the technique turns into money.

    Third, the full loop: how to turn a failed trace into an evaluation case. Step four feeds step two. Every failure that appears in production goes into the suite and never happens again without you finding out.

    Why it matters

    Why it matters. Today you know your agents run. You do not know what each appointment they produce costs, nor at which step they get stuck.

    The job advert asks for it in these words: tie agent actions to pipeline and revenue. Without that you cannot charge for outcomes, which is exactly the model you want to sell, nor defend a contract with a hospital that will ask you about unit costs.

    What you gain

    What you will be able to do that you cannot today. Say, with last week's data, what each agent did, what it cost and what it produced. Find the exact trace of a production failure in five minutes, instead of reconstructing it from memory. And set a price per outcome knowing your real cost per outcome.

    The deliverable: two Modal applications instrumented with traces, and one weekly line per agent in Slack. When that line exists and nobody has to assemble it by hand, the step is closed.

    Step 5 · Days 19 to 23 · 25 hours

    Reading and debugging Python

    Listen

    2 min 11 sDownload to listen offline

    Not learning to programme from scratch: learning to read fluently what Claude writes, and to spot the risk before it becomes an incident.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. The goal is reading and debugging, not writing from scratch. A more modest goal, and far more useful.
    2. Functions, lists and dictionaries, exceptions, modules and environments, files and JSON, API requests, environment variables and secrets, logging and tests.
    3. The three incidents that cost you most this year were all solvable by reading a few lines.
    4. You cannot review anyone's code, Claude's or a hired engineer's, without being able to read it.
    5. At six hours a day, CS50P weeks 0 to 7 fit into five days.

    Key concepts

    Traceback
    Python's error report. Read it from the bottom up: the last line says what happened, the ones above say where.
    Exception
    The error that interrupts the programme. Catching one without deciding what to do with it is the commonest way to hide a failure.
    Environment variable
    Configuration that lives outside the code. That is where your keys are, and where the switch that enabled sending from your main mailbox was.
    Unit test
    A piece of code asserting that another piece does what it claims. In Python, with pytest.

    Questions to revise

    1. Open one of your scripts at random: what does it do and how does it fail?
    2. Faced with a traceback, where do you start reading?
    3. Where do your secrets live today, and who can see them?
    4. Which of your 60 scripts has no test and would do the most damage if it got something wrong?
    Read the text you hear

    In one sentence

    Step five. Reading and debugging Python. Days nineteen to twenty-three, twenty-five hours.

    In one sentence, and this sentence matters: this is not learning to programme from scratch. It is learning to read fluently what Claude writes.

    The syllabus, and what is left out

    What you learn. Functions, lists and dictionaries, exceptions and what triggers them, modules and environments, files and JSON, requests to an API, environment variables and secrets, event logging and unit tests.

    And what is deliberately left out: algorithms, advanced data structures, object-oriented programming beyond being able to read it. None of that is needed for what you have to do.

    At six hours a day, weeks zero to seven of CS50P fit into five days. That is intense, but the material is designed for self-study and the problem sets are short.

    The three incidents of the year

    Why it matters. The three incidents that cost you the most time this year were all solvable by reading a few lines. The variable that enabled sending from your main mailbox. The local routine that never fired. The binary that was outside the PATH.

    None of the three was a hard problem. All three were expensive because the code was opaque to you. And there is a larger consequence: you cannot review someone else's code, neither Claude's nor that of the engineer you hire next year, without being able to read it. A reviewer who does not read is a rubber stamp.

    What you gain

    What you will be able to do that you cannot today. Open any of your sixty scripts and explain in two minutes what it does and how it fails. Read a traceback and know which line to look at. Write a pytest test unaided. And find a configuration risk by reading, before it turns into an incident.

    These five days are the hardest stretch of the accelerated calendar. If any part of the route needs a buffer day, it is this one.

    Step 6 · Days 24 to 27 · 22 hours

    Agents the Anthropic way

    Listen

    2 min 07 sDownload to listen offline

    The central technical piece of the job advert: the API from the inside, well-described tools, an MCP server with governed access, and the Agent SDK instead of a hand-written loop.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. Messages, system prompts, structured output, tool use, prompt caching and context budgets.
    2. A governed MCP server defines what it reads, what it writes, what needs approval and what is logged.
    3. The Agent SDK replaces the loop you currently hand-write in every Modal application.
    4. ghl-flow exists but has no scopes, no per-call logging and no write behind a gate.
    5. Rebuilding ghl-flow is your first portfolio piece as a go-to-market engineer.

    Key concepts

    MCP
    The protocol by which a model connects to external tools and data through a server you control.
    Scope
    The concrete permission a tool holds: what it may read, what it may write, and over which records.
    Approval gate
    The point where the agent leaves a proposal and a person decides. Without it there is no safe write into a CRM.
    Prompt caching
    Reusing the fixed part of the context between calls, to pay less and answer faster.
    Structured output
    Forcing the model to answer in a fixed format. Necessary, and never sufficient: a valid JSON does not constitute an approval.

    Questions to revise

    1. What does ghl-flow read, write and approve today?
    2. If an agent deletes a contact by mistake, where is that recorded?
    3. What would ghl-flow's single write tool be, and what gate would it have?
    4. How much of the loop you hand-wrote in Modal does the Agent SDK save you?
    Read the text you hear

    In one sentence

    Step six. Agents the Anthropic way: the API, tools, MCP and the SDK. Days twenty-four to twenty-seven, twenty-two hours.

    In one sentence: building the piece your technical portfolio is currently missing.

    The API from the inside

    What you learn. How the API works from the inside: messages, system prompts, structured output, tool use, prompt caching and context budgets. That set explains why one of your applications costs what it costs and why it sometimes answers slowly.

    How a tool is described so that an agent uses it well, which is step one applied with your hands.

    And how the Agent SDK saves you the loop: all that scaffolding of retries, limits and tool handling that you currently rewrite in every Modal application is already solved.

    A governed MCP server

    The heart of this step is learning to build an MCP server with governed access. Governed means four answers in writing: what it reads, what it writes, what requires approval and what gets logged.

    Why it matters. This is the central technical piece of the job advert, which puts it like this: MCP servers, skills and applications connected to the CRM. Your ghl-flow server exists, and that alone puts you ahead of almost any candidate. But today it has no defined scopes, no per-call logging, and no write tool behind a gate. In other words: it works, but it cannot be audited.

    What you gain

    What you will be able to do that you cannot today. Rebuild ghl-flow as a governed MCP server, with free reads, a single write tool that leaves a proposal for human approval, and every call logged. And explain it to a chief technology officer in ten minutes.

    That is your first portfolio piece as a go-to-market engineer. Not a course passed: a system another engineer can review.

    Step 7 · Days 28 to 30 · 18 hours

    Commercial systems, as an engineer

    Listen

    2 min 09 sDownload to listen offline

    You know GoHighLevel as an advanced user. The advert asks you to know it as an engineer: data model, shared definitions, enrichment with a cost per hit, and deliverability as a system.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. A CRM's data model from the inside: contacts, opportunities, stages, activities, fields and triggers.
    2. RevOps is mostly shared definitions: what counts as a lead, an appointment and a sale.
    3. Waterfall enrichment is measured by cost per hit, not by favourite vendor.
    4. Deliverability is a system: separate domains, warming, authentication and caps.
    5. The email engine has been paused since July over a problem a runbook solves.

    Key concepts

    Data dictionary
    The document fixing what each field and each stage means, so two people count the same thing.
    Waterfall enrichment
    Querying providers in order until the data appears, while measuring what each hit costs.
    Deliverability
    The probability that an email reaches the inbox. It depends on the domain, authentication, volume and complaints.
    24-hour window
    In WhatsApp, the period after the user's last message during which you may reply freely. Outside it you use a template and you pay.

    Questions to revise

    1. What exactly counts as an appointment at Flow, and does the whole team count it the same way?
    2. What does an enrichment hit cost you today?
    3. What has to be true before restarting the email engine without risking the main domain?
    4. Is the Score Cazador calculated by a script or by a person?
    5. How does Meta's WhatsApp price change of 1 October 2026 alter your cost per conversation?
    Read the text you hear

    In one sentence

    Step seven. Commercial systems, as an engineer. Days twenty-eight to thirty, eighteen hours.

    In one sentence: to stop using your commercial tools and start designing them.

    The CRM from the inside

    What you learn. The data model of a CRM from the inside: contacts, opportunities, stages, activities, fields and triggers. What actually happens when an automation moves an opportunity from one stage to the next.

    RevOps as a discipline, which sounds like jargon and is the opposite: it is shared definitions. What counts as a lead, what counts as an appointment, what counts as a sale, and everyone at Flow counting the same way. That is where half the scoreboards in this industry fall apart.

    Waterfall enrichment with a cost per hit: querying providers in order until the data appears, and knowing what each hit cost.

    Deliverability and WhatsApp

    And deliverability as a system, not as a trick: domains kept separate from the main one, warming, authentication, daily caps and watching complaints.

    Plus the WhatsApp rules: templates, the twenty-four hour window, and Meta's price change of the first of October, two thousand and twenty-six, which changes your cost per conversation.

    Why it matters. You know GoHighLevel as an advanced user; the advert asks you to know it as an engineer. And you have an email engine that has been paused since July because of a deliverability problem that willpower does not fix: a runbook does.

    What you gain

    What you will be able to do that you cannot today. Write Flow's revenue data dictionary, in a form another engineer can follow without asking you. Calculate the Score Cazador with a script instead of by hand. Restart the email engine without putting the main domain at risk.

    And run one complete circuit end to end, that is, list, agent, gate and appointment, measured on the scoreboard you built in step three.

    Step 8 · Days 31 to 33 · 15 hours

    Experimentation and propensity

    Listen

    2 min 27 sDownload to listen offline

    Designing a test before launching it, replacing a heuristic whose weights were set by eye with a model that learns, and not confusing correlation with effect in your own data.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. An A/B test is designed beforehand: written hypothesis, primary metric, sample size and stopping rule.
    2. Without a calculated sample size, the result is an impression with two decimal places.
    3. Logistic regression is the first propensity model, and it has one virtue: it explains itself.
    4. The Score Cazador is today a heuristic with weights set by eye.
    5. Even six weeks of scoreboard give you data to start replacing it.

    Key concepts

    Primary metric
    The only one that decides whether the test won. Chosen before seeing the data, and there is only one.
    Stopping rule
    When the test closes, defined in advance. Stopping when the number looks pretty is the classic way to fool yourself.
    Propensity
    The estimated probability that a contact does something, for example replies or books.
    Logistic regression
    The simplest propensity model. Its coefficients can be read and explained, which is exactly what the advert asks for.
    Effect against correlation
    That contacted doctors buy more does not prove contact works: you may have contacted the ones already about to buy.

    Questions to revise

    1. How many contacts do you need before starting a test, and how do you work it out?
    2. What is the primary metric of your next outreach test?
    3. Which variables do you think predict that a doctor replies, and which could you measure today?
    4. If a campaign coincides with a good month, how do you separate the campaign's effect?
    5. Why is a model that explains itself worth more than one that is slightly more accurate?
    Read the text you hear

    In one sentence

    Step eight. Experimentation and propensity. Days thirty-one to thirty-three, fifteen hours. It is the last numbered step and the shortest.

    In one sentence: swapping the impression for the result.

    Designing the test before launching it

    What you learn. The first thing is to design an A B test before launching it, not afterwards. That means four things written down: the hypothesis, the primary metric, the sample size and the stopping rule.

    The primary metric is one metric. If there are three, one of them always wins. And the stopping rule is fixed in advance, because halting a test on the day the number looks pretty is the most common, and most elegant, way of fooling yourself.

    You also learn to recognise noise. At small volumes, two identical sequences give different results. Knowing how many contacts you need before you start saves you months of false conclusions.

    Propensity that explains itself

    The second thing is propensity. Logistic regression as a first model: which variables predict that a contact will reply, and how to read the coefficients so you can explain the recommendation in words.

    That ability to explain itself matters more than a couple of points of accuracy. The advert asks for it literally: a predictive model that explains itself. A model that is accurate and cannot be explained can neither be sold to a doctor nor defended in front of a client.

    And the third: not confusing correlation with effect in your own data. That contacted doctors buy more does not prove the contact works, if you contacted the ones who were going to buy anyway.

    What you gain

    Why it matters. Your Score Cazador is today a heuristic with weights set by eye. It works, and it was the right call to get started. But even six weeks of scoreboard give you data to begin replacing it with something that learns and explains itself.

    What you will be able to do that you cannot today. Know whether one outreach sequence is better than another with a written result, win or lose. Rank the hundred daily outreach messages with a model that says why each one goes first. And show a client the effect of a campaign without crediting it with what would have happened anyway.

    Step 9 · Across days 24 to 33 · 8 hours

    A repository that looks after itself

    Listen

    1 min 42 sDownload to listen offline

    Getting the tests and the evaluation regression to run by themselves on every change, the main branch protected, and the skills versioned.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. Continuous integration: every proposed change triggers the tests without anyone remembering to run them.
    2. Protected main branch: nothing enters main without passing the evaluations.
    3. The step 2 evaluations become an automatic regression here.
    4. The audit of 11 September found skills with no version.
    5. Eight hours spread over the last ten days: one hour each afternoon, not a block.

    Key concepts

    Continuous integration
    A server that runs the tests on every proposed change and blocks whatever fails.
    Protected branch
    Configuration that prevents writing straight to main: everything goes through review and the tests.
    Regression
    Re-running the old tests to check that the new thing did not break what already worked.
    Skill versioning
    Each skill stating which version it is and what changed, so you can go back when something breaks.

    Questions to revise

    1. If Claude proposes a change to a skill today, what runs by itself before you see it?
    2. Which of your skills have no version?
    3. Can you say the sentence nothing reaches main without passing the evaluations, and have it be true?
    4. When a new engineer joins, what stops them if they get it wrong?
    Read the text you hear

    In one sentence

    Step nine, and it runs across the others: days twenty-four to thirty-three, eight hours in total, roughly one hour at the end of each of those days.

    In one sentence: making the repository defend itself, without you present.

    What gets automated

    What you learn. Getting the tests and the evaluation regression to run by themselves on every proposed change, without anyone having to remember. Getting the main branch protected, so nothing can be written to it directly. Review conventions, so that giving and asking for feedback does not depend on the mood of the day. And versioning for your skills, so you can go back.

    It is the shortest step and the least noticeable while you do it. It is also the one that saves the most work afterwards.

    Why it matters

    Why it matters. The advert asks you to grow a codebase with many contributors. Your audit of the eleventh of September found skills with no version, which is exactly the symptom of a repository that depends on one person's memory.

    And you are going to add engineers. The repository has to protect itself without you present, because if the only quality control is your personal review, you become the bottleneck of your own company.

    What you gain

    What you will be able to do that you cannot today. Receive a change from another person, or from Claude, and know within two minutes whether it broke something, because the tests have already run on their own.

    And be able to say the sentence nothing reaches main without passing the evaluations, and have it be true rather than an intention.

    Final revision · With a pause to answer

    Fifteen questions to revise

    Listen

    4 min 42 sDownload to listen offline

    The exam for the route. Fifteen questions read with a pause before each answer, so you can revise without a screen in front of you.

    Chapter index

    Tap a section to carry on from that point.

    The essentials

    1. Hear the question, pause, and answer before the answer arrives.
    2. If you miss three or more from one area, that step is not closed.
    3. The questions follow the order of the nine steps.
    4. At this pace, come back to this track every week, not only at the end.

    Questions to revise

    1. How many of the fifteen did you answer without hesitating?
    2. Which step did the ones you missed come from?
    3. Which deliverable of that step is still outstanding?
    Read the text you hear

    How this revision works

    Final revision. Fifteen questions across the nine steps.

    It works like this: you hear the question, there is a short pause, and then the answer arrives. Answer first yourself, out loud or in your head. If you miss three or more from the same area, that step is not closed, however many videos you have watched.

    Questions one to five: vocabulary and evaluations

    One. What is the difference between a workflow and an agent?

    A workflow follows steps you fixed in advance. An agent chooses its own next step, in a loop, until it finishes or runs out of budget.

    Two. Name three of the five agent design patterns.

    Chaining, routing, parallelisation, orchestrator with workers, evaluator with optimiser.

    Three. An agent is misusing a tool. What is the first suspect?

    The tool description. It is a prompt, not documentation.

    Four. How does an evaluation begin?

    By reading transcripts one at a time and noting in a single sentence what went wrong. Not with metrics.

    Five. How do you know whether your judge is any good?

    By measuring how often it agrees with the labels you applied by hand. First you evaluate the judge, then the agent.

    Questions six to ten: data, traces and code

    Six. What is Flow's fact table?

    The funnel: message, reply, appointment, proposal, sale. The dimensions are contact, campaign, agent, speciality and city.

    Seven. What is DuckDB for in this route?

    For querying your GoHighLevel exports, your Retell logs and your file of contacted doctors on your own Mac, without setting up a server.

    Eight. What is a trace, and what is a span?

    The trace is the complete run of an agent, seen as a tree. The span is each branch: a call to the model or to a tool, with its cost and its latency.

    Nine. What does a trace need before it can be joined to the scoreboard?

    The business tag: the contact and the campaign.

    Ten. In the Python step, what is the stated goal?

    To read and debug what Claude writes. Not to programme from scratch.

    Questions eleven to fifteen: MCP, commercial and method

    Eleven. Which four things define a governed MCP server?

    What it reads, what it writes, what requires approval and what gets logged.

    Twelve. What is RevOps, without the jargon?

    Shared definitions. Everyone counting a lead, an appointment and a sale the same way.

    Thirteen. Which four things are written down before launching an A B test?

    The hypothesis, the primary metric, the sample size and the stopping rule.

    Fourteen. Why prefer a propensity model that explains itself?

    Because you have to be able to tell a client why one contact goes first. A model that is accurate and cannot be explained cannot be defended.

    Fifteen, and the most important. What is the one non-negotiable rule of the free route?

    The hour spent reading transcripts from your own agents. At six hours a day, it is the first hour of every day. It is what replaces the paid cohort.

    Closing

    That is all of them. Twelve or more without hesitating and you are in good shape. If you missed several from the same area, go back to that track before carrying on with the calendar.

    And remember the criterion that runs through the whole route: during these thirty-three days, Claude reviews what you wrote first. It does not write it for you.