What should I measure in early-stage GTM experiments?
Measure whether one named assumption got stronger: the right people, an urgent problem, a next step that costs them something, and a first outcome they reach. Give each test one primary metric, one diagnostic, and a stop written before the first send. Keep acquisition-cost ratios and lifetime value off the kill rule until the same motion repeats.
Measure the assumption, not a growth total
Early-stage GTM experiments measure whether one named assumption got stronger or weaker: who the buyer is, whether the problem is current, whether one message earns a next step, or whether a new account reaches the promised outcome. Leads, clicks, meetings, and new logos are not that measurement.
Gustaf Alströmer, in YC Startup School, treats repeat use as the clean read. Registrations, visitors, conversion on unnamed people, net promoter scores, and surveys are weaker. A curve that keeps falling is a failing product. A curve that flattens is a retained set. Praise that exists only while the product is free is a bad sign when it should be paid.
Dan Hockenmaier, with Alströmer, calls heavy acquisition, especially paid, before any cohort levels off the usual early mistake. A modest activation gain usually beats a much larger jump in new users. See which GTM experiments to run first. If reached people never hit the outcome, read a product problem, not a GTM problem.
One primary metric and one diagnostic
Each experiment carries one primary metric, one diagnostic that explains it, and a decision you can say out loud.
Tracsio says a useful early metric shows whether one assumption strengthened, names the next change, and moves this week. Learning velocity is the loop: days to a written decision, assumptions the batch closed, whether it ended in a named next move, and how fast that move shipped.
| Experiment | Primary | Diagnostic | A keep means | It does not prove |
|---|---|---|---|---|
| Segment | ICP-fit calls naming a current problem | Unexpected titles, trigger | Next batch on this slice | Payment |
| Outbound | Positive reply rate on one list | Meetings from those replies, their words | Keep the angle | A scalable channel |
| Content | Qualified conversations with named readers | Topic and page | Keep the topic | That traffic was demand |
| Sales call | Next-step rate on ICP-fit calls | Urgency, workaround, who commits | Pursue the pain | Call quality |
| Offer | Scoped pilot or paid start | Cycle time, objection, restated value | Hold the offer | A market price |
| Onboarding | Activation inside a set window | Time to milestone, drop-off step | Fix that step | Later retention |
| Channel | Qualified conversations per hour, by channel | Whether they activate | More hours there | A blended cost |
| Pilot | Pilot-to-paid inside the window | Use, and whether the outcome happened | Keep that scope | A free pilot |
A positive reply shows curiosity or a willingness to continue. A neutral reply, a deck request or a circle-back, is not a yes. A refusal proves arrival, not fit. Stackmatix drops impressions, followers, and opens unless they connect to pipeline. Tracsio scores a call from 0 to 4 on unprompted pain, urgency, profile match, and a next step. Pain without urgency is timing. Urgency from the wrong role means the list is drifting.
Read the same four questions by segment
Every readout answers four questions, split by segment, channel, and use case. One blended rate hides the only slice that worked.
- Reach: ICP match among responders, not on a list you already filtered. Silence, and unexpected titles, are results. Validate the ICP first.
- Care: they state a current problem and name a workaround. No workaround is a question, not a kill. Urgency does not replace the next step.
- Act: a working session, a pilot with an end date, or a paid start. A meeting alone is a calendar entry.
- Value: activation inside the window, then retention on the natural cycle. Record why they left, or what they went back to, by segment. Expansion inside an account is a later line. Neither one is a cost ratio.
Alströmer says more channel work does not matter if people try once and leave. Pedowitz says to lock follow-up time, routing, and the talk track, and to use a matched group. If that follow-up slips, you measured the team. The buyer’s own time-to-reply is a diagnostic of urgency, not a primary metric, and not the same clock. Change one of audience, message, channel, or offer.
Activation is a timed milestone
Activation is the share of new accounts that reach a milestone of core value inside a window you wrote before they arrived.
Amplitude counts new users who hit core value inside a set window, not account creation. “Logged in” is too broad. A long advanced chain is too narrow. “Eventually” mixes a late accident with a first run. If hitters do not stay more often than people who miss it, change the definition.
Cohort by source and segment. If value needs weeks of hand-holding, Tracsio says the message is ahead of onboarding. Fix that path before you treat the product as the failure. Two milestones you can copy: the first real record in, not sample data, and a second person on the account. Alströmer sets the clock from the job: a pay cycle, irregular travel, or a daily habit. Day seven and day thirty fit only when a successful user would be back then. Hockenmaier treats a correlation with later retention as a clue, then tests the cause.
Hold cost ratios out of the kill rule
Hours per qualified conversation, and days from first contact to a decision, are usable while the customer count is still small. A fully loaded acquisition cost, payback, and lifetime value need a motion you can repeat and a margin you can see. One customer, paid or unpaid, is not a path you can run again.
Tracsio, citing David Skok, puts those ratios after a repeatable process. Early revenue does not show whether the break is the message, the list, the offer, or the first run. UpliftGTM says scale metrics before launch create false confidence.
The Stackmatix seed note names ICP validation, trial-to-paid, and cost by channel. Its launch table already includes payback, and its later stage jumps to lifetime ratios and a seven-day activation clock. Do not take that jump. Add payback only after the same motion repeats, state the account count, and do not blend channels. See when a seed motion is ready to scale. Price is a start date, not “would you pay.”
How long should a GTM experiment run before you decide?
Run the shortest window that can answer the hypothesis, and write the end date and the minimum count before the first send. Tracsio says outbound often shows a direction in days, 7 to 14 days is often enough for a message, and content more often needs 3 to 4 weeks: neither number is a minimum.
Tracsio says to stop when a large enough sample contradicts the hypothesis. Extend when the count is thin or the direction is improving. Do not keep the test alive on one encouraging reply. Do not move the date after a weak start. Split a mixed result before you kill it. Pedowitz puts many meeting pilots at 2 to 6 weeks, and longer when the read is win rate.
The GTM Playbook wants two full weeks so one weekday cannot decide, and calls a classic test premature with no baseline. At seed the first precommitted batch is that baseline: keep, change, or stop for the next batch only. A mid-window edit restarts the count. Hockenmaier notes a stark contrast needs less volume than a wording tweak. Friends and one logo fail the audience gate.
Close the log before you open the next test
You have learned from a test when a stranger can read the hypothesis, the rule, the result, and the sentence that changes the next test.
The GTM Playbook keeps id, dates, category, hypothesis, baseline, decision rule, result, decision, and a one-to-three sentence learning note. “The old line won” is not a learning. Before the send, write the pass, the inconclusive band, and the miss.
Paste quotes, the workaround, and the objection by segment, and note any second pitch or price change in the window. If the buyer is in the EU, a quote is personal data: keep it for a reason, and delete it when the test closes. Do not track opens: a weak metric, and in the EU a consent problem. The GTM Playbook sets a 30 to 45 minute weekly review: one owner, one signal and one recommended move per input. On a one-person team, that owner writes the line. Tracsio ends the review on one move: keep the hypothesis and run it again, narrow the slice, change the message, fix the drop-off step, or stop. Build a workflow only after the same card holds on a second batch.
Can AI keep the log without picking the winner?
AI can maintain the log and draft the readout once every field has one definition. It should not choose the metric, invent a missing baseline, or declare the winner.
Sort replies only after a person has labeled a sample. An empty cell stays empty. Automating the report drafts the note and must not invent the total. The model does not merge segments, move the date, or rewrite the stop. Hockenmaier puts queryable events ahead of any experiment tool.
Cut a metric that cannot change the next move
If you cannot name the action when a number rises or falls, take it off the weekly card.
Alströmer tells early teams to skip the split test. A significance calculator belongs to a later test that already has volume. Leave the card unbuilt while qualified, positive reply, or activation is undefined, while the offer changes every call, or while nobody reaches the outcome.
The scorecard I would leave in your accounts
One hypothesis, one variable, a primary metric, a diagnostic, a count, a date, and a stop.
I am one person, under the name Poldermarketing, remote, in Dutch and English. Hire me freelance for a bounded project or a few days a week. I am an AI-native marketer, strong in AI, content, automation, and building workflows, and I build and run the work rather than only advising. Google Ads and Meta Ads are relatively new to me: I set them up and review them, and I do not scale a large media program.
That is one marketer for marketing, AI, and automation for a freshly funded startup, fully remote, from the first message to the first customers, without a separate specialist for every part. The free growth scan shows how the site reads. How I work is the engagement. Positioning and messaging is where I start when the sentence is the constraint.
Questions people ask
Should acquisition cost and lifetime value be on the first scorecard?
No. Acquisition cost and lifetime value need a motion you can run again and a margin you can actually see. On the first batches, track hours per qualified conversation, the next-step rate, and whether new accounts hit the milestone you defined. Add the cost ratio only after a second batch repeats the same path. One unpaid introduction will make any ratio look kinder than the motion is.
Is a high reply rate a successful experiment?
Only when the replies are positive, from the segment you named, and they lead to a next step. A refusal still proves the message arrived. It does not prove urgency. Keep the rate, and read the diagnostic next to it: meetings from those replies, the phrasing buyers used, and whether those accounts later reach the milestone. If they never do, the test showed you can get attention, not that you have a motion.
What if two segments show opposite results?
Do not average them. A blended rate hides the only slice that worked and makes a real signal look weak. Split the primary metric and the diagnostic by segment before you choose. Keep the slice where buyers named a current problem and took a next step. Stop or rewrite the slice that stayed silent. A job title you did not expect, repeating with real urgency, becomes its own row on the next test.
Do I need statistics software before I can stop a test?
No. Early volume will not support a textbook split test, and waiting for one is how a test never ends. You need a count and an end date chosen beforehand, a batch big enough that one friendly reply cannot rescue it, and a written decision when you hit that gate. A significance calculator belongs to a later test that already has a baseline. It is not a substitute for the stop rule.
Should every experiment track day seven and day thirty retention?
Use those checkpoints only when a successful customer would naturally return on that clock. A daily product can use a short window. A payroll or other monthly product will look failed if you insist on a weekly return. Pick the window from the job, write it into the test before anyone signs up, and read it by source and by segment. Dashboard defaults are not a definition of value.