Back to all articles

pr

How to Audit a Pitch With 37 Claude Subagents

A practical method for pressure-testing pitches, press releases, and launch announcements with isolated Claude subagents, plus the limits of synthetic personas.

Elvis SunElvis Sun
View as MarkdownHow to Audit a Pitch With 37 Claude Subagents

Before a pitch, press release, or launch announcement reaches the public, you can give it 30 isolated first reads and combine them into one ranked audit. The important word is isolated: 30 personas inside one Claude chat are not 30 independent opinions; they are one model continuing one context.

I tested the difference on the Medialyst landing page. One prompt fanned out into 37 Claude subagents: 30 persona reactions, six segment-level syntheses, and one final cross-segment verdict. I ran the identical prompt in a normal Claude chat beside it. That control is the part worth understanding, because it shows why the architecture matters more than the persona names.

Watch the full 11:32 walkthrough on YouTube. It includes the side-by-side control run and the complete workflow described below.

What problem does a subagent audit solve for PR teams?

Comms teams routinely have to publish material before they can see how every relevant audience will read it. A pitch might make sense to an in-house comms lead and sound vague to a journalist. A launch announcement might persuade a founder but leave a marketer unsure what the product actually does.

The tempting shortcut is to paste the draft into one AI chat and ask it to react as 30 different people. The output looks like a research panel. It is not one.

In a single context, the model generates persona two after persona one, persona three after both, and so on. Later responses are conditioned on what came earlier. Instead of 30 first impressions, you get one long sequence wearing 30 name tags.

A subagent audit fixes that specific problem. Each persona reads the artifact in a separate context, writes a reaction without seeing the others, and only then feeds into synthesis. For PR teams, the method is useful as a pre-screen: find obvious confusion before the pitch reaches a journalist or the release goes on the wire.

Why can't one Claude chat give you 30 independent opinions?

Imagine running a focus group by asking 30 people to speak one at a time, while everyone else listens. The first person makes a confident point. The next person now has that point in mind. By person 20, the conversation may be responding more to the room than to the material you wanted tested.

That is roughly what happens inside one context window. The chat does not reset when the label changes from “PR agency owner” to “in-house comms lead.” Every previous token remains part of the material conditioning the next answer.

You can think of it as anchoring bias for token prediction. A later persona may agree, disagree, or try to add novelty, but it is still generated in the shadow of the earlier personas. The output has plurality in format, not independence in method.

The practical rule is simple: if you want separate first reactions, separation has to happen before synthesis. Do not build 30 characters in one room. Build 30 rooms.

What counts as an independent subagent reaction?

For this kind of audit, an independent reaction has four properties:

  1. One persona gets one isolated context.
  2. Every persona receives the same target artifact and the same audit brief.
  3. No persona can read another persona's output.
  4. Reactions are combined only after the individual runs finish.

This is context independence, not statistical independence and not proof that the personas resemble real buyers. All 30 are still generated by the same underlying model. Isolation removes the most obvious cross-persona contamination; it does not turn a model's guesses into human research.

That distinction is the foundation for using the method honestly.

How did the 37-agent landing-page audit work?

I used Claude Opus 5 and pushed effort to the maximum setting before sending the prompt. In this run, the result was a parallel fan-out rather than one agent performing every voice sequentially. That is an observation from one run, not evidence that maximum effort alone caused the behavior or that the setting is a hidden feature-unlock switch.

The audit used one prompt and three phases:

PhaseAgentsJob
Independent reactions30Each persona read the live Medialyst landing page in its own context and wrote a reaction.
Segment synthesis6One synthesizer per audience segment read the five reactions in that segment.
Cross-segment verdict1One final agent read the six segment summaries and produced the overall audit.

The count matters less than the boundaries. Phase one maximized separation. Phase two allowed comparison within a relevant group. Phase three looked across groups only after the independent work had survived intact.

The final output contained five ranked fixes, a score for each persona across several categories, and a section on what would win over each segment. Those deliverables made the reactions easier to compare, but the underlying benefit came from the fan-out, not the scorecard. The five fixes are not disclosed in the source material for this article, so the reproducible evidence here is the architecture and run receipt, not the substance of those withheld findings.

What exact prompt launched the 37-agent audit?

This was the complete input, reproduced with the original agents to testing typo:

look at my site https://medialyst.ai

and spawn 30 agents to testing my messaging and positioning. 5 each with slightly different persona within the group:
- PR agency owner
- in-house comms team
- PR freelancers
- founder
- marketer
- journalists

I pasted it into Claude Code, running locally on a worktree branch, after raising the effort selector. I did not provide a skill, orchestration script, or config file. Claude wrote the orchestration and ran the subagents from that one prompt.

Which perspectives were included in the audit?

The 30 personas were divided into six groups of five:

  • PR agency owners
  • In-house comms leads
  • Freelancers
  • Founders
  • Marketers
  • Journalists

The roster included solo founders, a Head of Comms, a VP Comms, an ex-agency freelancer, fractional CMOs, a technical solo founder, a Series A founder, a non-technical DTC consumer-brand founder, a tech reporter at a major outlet, and an editor at a B2B trade publication.

That range gave the audit different lenses on the same page. It did not make the panel representative of a market. Persona variety broadens the pre-screen; it does not validate the persona assumptions.

For a pitch or press release, the roster should follow the artifact. A journalist-facing pitch needs reporter perspectives. An internal announcement may need employees and managers. The reusable part is not this exact list. It is the combination of deliberate segments and isolated contexts.

What did the normal-chat control run reveal?

I pasted the identical prompt into a normal Claude chat and ran it beside the subagent version. The normal chat produced roughly one line per persona, and I abandoned it partway through.

The problem was not only brevity. Every later persona was being generated after the earlier persona's answer. The control put all 30 simulated people in the same room and asked them to take turns.

The subagent run changed the information flow. Each of the 30 persona agents reacted to the landing page, not to a growing transcript of other reactions. Only the six synthesizers were allowed to compare outputs, and comparison was their explicit job.

This is why a long role-play prompt is not a cheaper version of the same workflow. It is a different workflow with a different failure mode.

This comparison was one worked example, and the normal-chat control was abandoned partway. It shows the difference in information flow by construction; it does not prove that isolated subagents produce materially better findings across repeated runs. A stronger comparison would repeat both setups and measure issue diversity, overlap, and stability. That experiment was not part of this run.

How long did the audit take, and what did it cost?

The run took roughly 20 minutes. The last on-screen token count was 2.3 million while a phase was still running, so that figure is a mid-run reading rather than a final total.

Using the supplied Opus 5 prices, the simple full-rate boundaries on that visible 2.3-million-token count are $11.50 if every token were input and $57.50 if every token were output. The stated estimate for this run is roughly $15–$25 because the fan-out was heavily input-weighted and cache reads bill at a fraction of the full input rate.

That estimate is not a reconstructable invoice total. The source does not contain the final token count, the input/output split, or the cache-read volume, and the visible count was captured before the run finished. Treat $15–$25 as a run-shape estimate, not precise accounting.

Traditional message testing and consumer research can enter five figures and take weeks. That makes the subagent audit attractive, but the comparison needs a bright line: the outputs are not equivalent.

The Claude run is a same-afternoon pre-screen with synthetic personas. Research with real participants is how you investigate the questions a model cannot answer. The useful economic argument is not “replace the study.” It is “do not spend research budget discovering the copy was obviously confusing.”

What can synthetic personas actually tell you?

Synthetic personas can surface parts of an artifact that appear confusing, unsupported, or unconvincing from several modeled perspectives. They can help a team notice repeated themes, compare how the same language lands across segments, and produce a shortlist of issues to inspect.

In this run, the useful deliverable was a ranked set of five fixes plus segment-level guidance. The ranking turns a large pile of reactions into an editing queue. The scores make the reactions sortable. The “what would win them over” section suggests questions for the next revision or the next research step.

None of those outputs becomes true because 30 agents generated it. Treat repeated findings as hypotheses with stronger internal support, not facts about a market.

What can't synthetic personas tell you?

A synthetic persona is Claude's guess at a buyer, not a buyer. It cannot tell you what someone will actually pay for, whether a journalist will cover the story, or how a real customer will behave after reading the page.

NN/g's published position on synthetic users is skeptical. That skepticism is the right baseline here. The method becomes more credible when its claim is narrower: it is a way to inspect messaging from multiple isolated model-generated perspectives before involving real people.

Context isolation fixes contamination, not syntheticness.

Use the audit to find the obvious problems cheaply. Then spend interviews, message-testing, or consumer-research budget on the non-obvious questions. In other words, it does not replace a study. It helps you decide whether you need one and where to point it.

How can you run the same pre-screen on a pitch or press release?

The workflow transfers cleanly to any artifact where several first reactions would be useful before publication:

  1. Choose one artifact. Start with a pitch, press release, launch announcement, landing page, onboarding flow, or spec. Keep the target stable across every persona.
  2. Define the audience segments. Pick groups that have genuinely different relationships to the artifact. Then create several concrete personas inside each segment.
  3. Fan out the independent reads. Give every persona the same artifact and evaluation questions, but run each in its own context. Ask for consistent outputs such as category scores, the main points of confusion, and what would make the material more convincing.
  4. Synthesize within segments. Once all independent reactions are complete, give each segment's responses to a dedicated synthesizer. Its job is to find agreement, disagreement, and priorities inside that group.
  5. Synthesize across segments. Give the segment summaries to one final agent. Ask it to rank the cross-cutting issues and preserve meaningful disagreements rather than flattening them.
  6. Separate edits from research questions. Fix obvious clarity problems. Turn claims about purchase intent, journalist interest, or real behavior into questions for real people.

The sequence is the method. If synthesis starts inside the persona-generation context, the audit loses the independence it was designed to create.

What should a pitch or press release audit ask?

The landing-page run demonstrates the orchestration pattern, not a PR-specific result. For a pitch or release, turn the audit brief into questions the team can act on:

  1. What is the news value, in the recipient's own terms?
  2. Why is this relevant to the target journalist or audience?
  3. Which claims are substantiated by the material provided?
  4. What reputational downside or likely objection needs human review?
  5. What proof is missing?
  6. What action is the recipient being asked to take?

Keep those questions identical across personas. The point is to compare isolated reactions to the same brief, not let every agent invent a different test.

What is a smaller five-person starter version?

Thirty-seven agents were the worked example, not a minimum. A smaller team can preserve the key boundary with five isolated reactions and one synthesis. Paste a top-level instruction like this into Claude Code and replace the bracketed fields:

look at [artifact or URL]

spawn 5 agents to test it, each with a different [audience segment] persona.
Run every reaction in an isolated context. Each agent should return:
- the main message it understood
- what was confusing or unconvincing
- category scores using the same categories
- what would win this persona over

after all 5 finish, spawn 1 synthesizer to report:
- agreement and disagreement
- edit now
- validate with the client
- test with real people

That is a starter adaptation, not the verbatim prompt used in the 37-agent run. The original run used maximum effort and produced parallel fan-out; this single example does not establish which settings are necessary in every run.

How should an agency keep the workflow reviewable?

The agent architecture does not decide who owns the result. Add a human operating layer around it:

  1. Name one owner who approves the persona definitions and resolves conflicting recommendations.
  2. Before fan-out, confirm that the draft is approved for use with the chosen model under the team's material-handling rules.
  3. After synthesis, discard duplicate or unsupported model speculation, then route the rest into three buckets: edit now, validate with the client, and test with real people.
  4. Require an approval checkpoint before any change reaches the pitch, release, wire, or journalist.

This keeps scores and ranked fixes from becoming an automatic to-do list. They remain inputs to a comms decision.

How should you use the final verdict?

Use it as a map of hypotheses, not a vote count.

If the same confusing passage appears across several isolated reactions and survives segment synthesis, it is a reasonable candidate for revision. If a concern appears in only one segment, that may be exactly what makes it important, but it deserves investigation rather than automatic acceptance.

Treat persona scores as an organizing device. They let you sort and compare reactions; they do not measure market truth. The highest-value output is usually not the final number. It is a clearer answer to two questions: what can we fix now, and what do we need to ask a real person?

That is the useful place for Claude subagents in message testing. They make a cheap first pass more rigorous by keeping the first reads separate. They do not make synthetic users real.

FAQ

Are Claude subagents a replacement for a focus group?

No. The subagents are model-generated personas, not real participants. Use them to pre-screen messaging and surface hypotheses, then use real research for willingness to pay, behavior, and other questions that require actual people.

Why not ask one Claude chat to role-play every persona?

Because each later persona is generated with the earlier answers still in context. That creates anchoring and cross-persona contamination. Separate subagent contexts preserve independent first reactions until the synthesis stage.

Does maximum effort unlock Claude subagents?

No. I set effort to maximum and then observed parallel fan-out in this run. One run does not establish that the setting caused the fan-out or is required in every case, and it should not be described as a hidden switch that unlocks the capability.

How much did the 37-agent audit cost and how long did it take?

It ran for roughly 20 minutes. The last visible token figure was 2.3 million mid-run, and the estimated compute cost was about $15–$25 using the supplied Opus 5 prices and the run's input-heavy, cache-assisted shape. The exact bill cannot be reconstructed without the final count, input/output split, and cache-read volume.

What PR materials can this method audit?

Use it on a pitch before it reaches a journalist, a press release before it goes on the wire, or a launch announcement before publication. The same structure also works for landing pages, onboarding flows, and specs.

Find the right journalists — instantly.

Medialyst builds you a vetted media list in seconds. Just describe your story.

Try it free