What you'll be able to do after this guide: add logging to every agent so you can see where prospects are lost, tell a quiet day from a broken agent, and spot bugs before they cost you prospects.
The funnel log is a single CSV file (a plain spreadsheet file any tool can open) shared by every agent. Each run appends one row per step:
| run_date | run_ts | task | phase | step | count | detail |
|---|---|---|---|---|---|---|
| 2025-05-12 | 07:51 | 1 - Source & Research (Segment A) | in | search_candidates | 60 | |
| 2025-05-12 | 07:51 | 1 - Source & Research (Segment A) | filter | deduped_existing | 38 | |
| 2025-05-12 | 07:51 | 1 - Source & Research (Segment A) | filter | status_disqualified | 3 | 2 dissolved 1 in administration |
| 2025-05-12 | 07:51 | 1 - Source & Research (Segment A) | filter | unverifiable | 4 | |
| 2025-05-12 | 07:51 | 1 - Source & Research (Segment A) | out | researched | 15 | |
| 2025-05-12 | 07:51 | 1 - Source & Research (Segment A) | out | reached_review | 6 |
The human log is a monthly text file. Each run adds a heading and one line that reads like a funnel, plus one sentence on anything unusual:
1 - Source & Research (Segment A), 12 May 07:51. 60 in, minus 38 already in CRM, minus 3 not active, minus 4 unverifiable, 15 researched, 6 to Review. Two companies' accounts were scanned images and needed text recognition.
The CSV is for adding up; the text file is the one people actually read.
Filter rows record how many were removed, not how many survived. Removals tell you where the bottleneck is. If "deduped_existing" climbs towards "search_candidates", your search keeps returning companies you already have, and it's time to widen it.
Use a shared list of step names. Keep the allowed step names in one document. If one agent writes "dedup" and another "deduped_existing", the numbers can't be compared. Adding a step means updating the list in the same edit.
Log what happened, not what should have happened. If a step was skipped because a tool was down, log it with the reason in the detail column. A missing row looks the same as a genuine zero.
A run that did nothing still logs. Otherwise an empty run and a run that never started look identical. For several agents, an empty run is the normal, healthy result.
Logging never changes an outcome. If the log can't be written, the agent says so in its summary and carries on. A lost log line is cheap; a lost send or a duplicated task isn't.
Keep the detail column comma-free, or the CSV breaks.
Write a "what to watch" line into each spec:
| Agent | Watch | What it tells you |
|---|---|---|
| Sourcing | Share of candidates removed as already in the CRM | When your search is saturated |
| Backfill | Share of records scored as weak fits | Whether hand-added and imported records are worth having |
| First touch | Share of candidates already actioned | Whether sourcing volume is the real constraint |
| Acceptance check | Accepted divided by requests sent | Your acceptance rate, the most useful number in the pipeline |
| Email task | Tasks not created because of the cap | Whether one person can keep up |
| Deal sync | Replies divided by conversations checked | Your reply rate |
Some counts must always line up. If they don't, something is broken:
Deal comments written should equal messages sent where a deal existed.
Deals moved to Contacted can't exceed deal comments written, because the comment is written first.
An agent that should never route partners to the social channel should log zero for that step. Anything else means something upstream changed.
Agents sometimes skip logging. Row counts tell you the agent ran at least that often, never exactly how often. The scheduler is the only reliable record of when tasks ran.
Every agent's last reply is a plain summary for a person. Beyond the counts, ask it to name, explicitly:
Any write that failed, with the record.
Any tool that was unavailable and what it affected.
Anything that contradicted the instructions.
Any judgement call worth reviewing (a candidate close to an exclusion line, a company it chose to leave at Researching).
Any decision it is waiting on from a person.
When you put off a decision about the pipeline, ask the agent to report the behaviour it affects in every summary. A deferred decision should stay visible.
Next: Steer Without Rewriting Prompts.