Prompt Engineering 10 min read Updated Jul 14, 2026

Set Kill Criteria for an Option With AI

You ask AI "should we keep pursuing this?" and get "it could be valuable — keep testing," so a weak option survives on hope. Here's how to set kill criteria first: the observable evidence, threshold, deadline, and owner that decide, in advance, when to drop it.

Build a Kill-Criteria Prompt

When "it could be valuable" keeps an option alive forever

You ask the AI whether to keep pursuing something — a feature, a vendor, a content angle, an experiment — and it answers the way it always does: "It could be valuable, but there are risks. Consider testing it further before deciding." Every word is reasonable, and none of it helps. There's no line that would tell you the option has failed, no date by which you'd expect an answer, no cost you're not willing to exceed. So the option survives — not because it's working, but because "could be valuable" is always true of something, and nothing was ever set that would let you say it's time to stop. Three weeks later you're still testing it, a little more attached, a little more spent.

That's the gap this guide closes. Evaluating an option isn't only "is this good?" — the harder, more useful question is "what would make us stop?", and it's far easier to answer before you're invested than after. Kill criteria are the abandon rules you write up front: the observable evidence, thresholds, and deadlines that, if they hit, mean this option no longer deserves to continue. They aren't a decision matrix scoring choices against each other, and they aren't a vague "if it doesn't work out" — they're specific, measurable triggers agreed before the sunk cost accumulates. NewPrompt helps you draft those criteria; it doesn't track a metric, watch a deadline, or retire an option for you. The model proposes the triggers when you run the prompt in your own AI tool, and whether one has actually fired — and whether that means stop — is the owner's call, not the model's.

Why AI keeps every option on life support

The model isn't hedging to be evasive — it's doing what a balanced answer looks like, and a balanced answer rarely says "quit." Asked to evaluate, it lists strengths and weaknesses and lands on "keep going, carefully," because that reads as thoughtful and offends no one. But a recommendation with no stopping line is a recommendation to continue by default. Four reasons a weak option outlives its usefulness:

  • "Promising" has no opposite. "It could work" is unfalsifiable — there's no result that contradicts it, so the option can never fail the test, because there is no test.
  • No threshold, no failure. Without a number or a signal that counts as "not good enough," every result gets read as "needs a bit more data," and more data is always available.
  • No deadline, no reckoning. An option with no date to decide by is never overdue, so the review that would end it keeps getting deferred to a next week that doesn't arrive.
  • Sunk cost stays invisible. The more you've put in, the more "let's not waste that" argues for continuing — and nothing in a strengths-and-weaknesses answer makes that pull visible, so it wins quietly.

Step 1: Name the assumption each criterion is testing

A kill criterion only makes sense against a belief it's checking, so start by making the assumptions the option rests on explicit. Whatever you're pursuing, it's alive because you believe a few things are true: that you can build it in the time you have, that users want it, that it won't derail something more important, that it's accurate or safe enough to ship. Write those beliefs down as plain statements, because each one is a place the option can fail — and a kill criterion is just an assumption paired with the evidence that would prove it false.

The reason to lead with assumptions rather than jump to "what's the metric" is that it stops you from measuring the easy thing instead of the deciding thing. An option usually dies on the assumption nobody wrote down — "we assumed this wouldn't delay the core work" — not on the metric that was convenient to track. Ask the model to surface the assumptions the option depends on, including the uncomfortable ones, so the criteria you set next attach to the beliefs that actually carry the decision, not just the ones that are simple to count.

Step 2: Turn each assumption into an observable trigger

Now convert each assumption into a criterion you could actually watch fire — the difference between "if it doesn't work" and "if fewer than 2 of 5 target users call it must-have." A usable kill criterion names the evidence (what you'd look at), the threshold (the specific value or signal that counts as failure), and the source (where the number comes from). "Too slow" isn't a criterion; "the basic version takes more than two engineer-weeks, by the tech lead's estimate" is. The test is whether two people looking at the same evidence would agree the trigger has hit — if they could argue about it, the criterion is still too vague to end anything.

This is the discipline of deciding the bar before the data arrives, which is what keeps a failing result from being rationalized into a passing one. The Product Validation Measurement Plan Prompt is a worked model of it: it turns a success metric into the specific signals and cuts that prove or disprove it, and — the part that matters here — it fixes what counts as hit, partial, or miss before the numbers are in, so the verdict can't be moved to fit what you hoped. A kill criterion is the same move aimed at the exit instead of the win: the threshold that means stop, written while you can still set it honestly.

Step 3: Attach a deadline, a data source, and an owner

A trigger with no clock never fires, so every criterion needs a date: the point by which the evidence should exist and the check will happen — "before sprint planning," "after five user interviews," "at the prototype review." The deadline is what turns a criterion from a hope into a commitment, because it names the moment the option has to show its evidence or forfeit the benefit of the doubt. Without it, "we'll know when we see it" becomes a review that never gets scheduled.

Each criterion also needs a data source and an owner — where the evidence comes from, and who is responsible for checking it and raising the flag. The owner matters because a criterion nobody owns is a criterion nobody watches: the tech lead owns the build-effort trigger, the PM owns the user-demand one, the person who can see the core delivery plan owns the does-this-delay-the-MVP one. Assigning it makes the check somebody's job on a specific date, rather than a good intention that dissolves into the general hope that someone's keeping an eye on things.

Step 4: Decide the action if it triggers — and route it to a human

A criterion that hits should point to an action, not just a feeling. For each one, write what happens if it triggers: kill the option outright, defer it to a later phase, cut its scope to a cheaper version, or escalate the call to a decision-maker. "If the build estimate exceeds two engineer-weeks, move it out of the MVP and revisit next quarter" is actionable; "if it's too expensive, reconsider" is not. And leave room for an honest exception — a review note that lets an owner override a triggered criterion with a stated reason, because a pre-committed rule should inform the decision, not replace the judgment of the person accountable for it.

One line is non-negotiable: the criterion informs the decision; it doesn't make it. When a trigger fires, the model doesn't retire the option and neither does the criterion on its own — the owner or the stakeholders look at what fired, weigh it against everything they now know, and decide to stop, continue, or change course. The point of writing the criteria in advance isn't to automate the kill; it's to make the moment of decision honest, so "we're stopping because we agreed this signal meant stop" is available instead of a debate driven by whoever is most attached. The Product Validation Decision Framework Prompt models that final read — weighing the evidence into a clear verdict and an iterate, pivot, or scale call, with the open risks that could change it — the deliberate decision a fired criterion should trigger, not skip.

Step 5: Get the criteria approved before the option proceeds

Kill criteria only work if they're agreed before the pursuit starts, so the last step is sign-off, not filing. Take the drafted criteria to the owner and the stakeholders and get them agreed while everyone can still be objective — because the entire value of a pre-committed trigger is that it was set when no one was attached, and a criterion written after the results are already known is a post-hoc rationalization wearing a rule's clothes. If you must add or change a criterion mid-flight, label it as such and say why, so the record stays honest about what was decided when.

Then put the criteria somewhere they'll actually be checked, and read the whole set for what it is. NewPrompt drafts the criteria; it doesn't monitor your metrics, watch the deadline, send a reminder, or retire the option when a trigger fires — the criteria go into your own tracker, calendar, or review process, and the checking is a person's job on the dates you set. The model can also miss a criterion that would have mattered or propose a threshold that's wrong for your context, so the draft is a starting point the owner sharpens, not a verdict. Set the exits before you're attached, and the hard decision later gets easier; whether to actually take one when a trigger fires stays with the people who own the option and its consequences.

Common mistakes

The habits that let a weak option outlive its evidence:

  • Setting only success criteria, never kill criteria. Knowing what winning looks like doesn't tell you when to quit; write the abandon triggers too, and write them first.
  • Writing vague triggers. "If it doesn't work" never fires because it's never clearly true; use an observable evidence, a threshold, and a source two people would read the same way.
  • Leaving out the deadline. A criterion with no date is never overdue, so the review that would end the option keeps getting pushed; name the moment the evidence is due.
  • Setting criteria after the results are in. A threshold written once you know the outcome is a rationalization, not a rule; commit the criteria before the pursuit, and label any later change.
  • Assuming the trigger auto-kills. A fired criterion informs a decision; the owner still weighs it and decides — stop, continue, or rescope — with room for a stated exception.
  • Filing the criteria and forgetting them. NewPrompt doesn't track them; put the deadlines and owners into your own process, or the criteria never get checked.

A worked example: AI chart generation in an MVP

Take a feature that would otherwise survive on "could be valuable," and set the criteria that would make you cut it — before the team is attached to it.

A "should we keep it?" question returns a vague "keep testing"; pre-committed kill criteria give each assumption an observable trigger, a deadline, an owner, and an action — decided before the team is attached
THE OPTION:
  Include AI chart generation in the MVP (plain-language -> charts).
  Constraints: MVP ships in 6 weeks; the core workspace and manual
  dashboards are higher priority; we cannot ship misleading analytics;
  engineering capacity is limited.

THE WEAK ASK, AND WHAT IT RETURNS:
  ask: "Should we keep AI chart generation in the MVP?"
  -> "It could be valuable but risky. Consider testing it with users
     before deciding."  (no threshold, no date, no owner -- it survives)

A BETTER ASK (kill criteria, decided before we commit):
  "Set kill criteria for this option before we continue. For each
   assumption give: kill_criterion | evidence | threshold | deadline |
   owner | action_if_triggered. Criteria must be observable; include a
   cost/time cap and an accuracy trigger; do not recommend a final kill
   automatically -- route it to the owner."

THE KILL-CRITERIA TABLE IT DRAFTS:
  assumption          trigger / threshold        deadline    owner   action
  build is cheap      estimate > 2 engineer-     sprint      tech    move out of MVP,
                      weeks                       planning    lead    revisit next qtr
  won't delay core    any core-workspace         weekly      product kill or defer
                      milestone slips for it      check       owner   the feature
  accurate enough     > 10% of charts in a test   prototype   data/   do not ship
                      set are materially          review      product
                      misleading
  users need it now   < 2 of 5 target users call  discovery   PM      remove from MVP
                      it MVP-critical             review
  can't mislead users ships without a guardrail   before      design/ block until the
                      that flags misleading or     beta        QA lead guardrail ships
                      low-confidence charts

HOW IT'S USED:
  each trigger has a date and an owner; if one fires, that owner brings
  it to the team, and the team decides stop / defer / rescope. The model
  didn't kill the feature -- it wrote the lines that make the later
  decision honest instead of driven by whoever built it.

STILL YOUR JOB:
  approve the criteria before the work starts; put the dates in your own
  tracker; NewPrompt doesn't watch the metric or fire the trigger.

Where this fits in NewPrompt

Setting kill criteria is a move you make on one option, and the pieces are small. The Markdown Output Builder holds the criteria in a fixed table — assumption, trigger, threshold, deadline, owner, action — so each one comes back with its evidence and its exit attached; the System Prompt Generator can bake "set kill criteria before committing to an option" into a standing instruction so the model volunteers the exits instead of only the upside. Two resources show the same before-the-data discipline: the Product Validation Measurement Plan Prompt fixes hit, partial, and miss thresholds before the numbers arrive, and the Product Validation Decision Framework Prompt is the deliberate read a fired trigger should lead to. They structure the thinking; the watching and the deciding are yours.

It's worth placing this against the decision moves it sits near. Comparing options with a decision matrix picks the best of several; deciding when to defer asks whether to make a call now or wait; classifying a decision as reversible tells you how much undo you have; asking what would change an answer maps the conditions a recommendation rests on. Kill criteria run alongside a chosen or in-progress option and answer a different question than all of them: not which to pick or when to decide, but what evidence, agreed in advance, would make us stop pursuing this one. Reversibility feeds in as one trigger among several — "if we'd lose the ability to back out, that's a stop" — but the frame is abandonment, not selection.

Because the expensive failure is rarely the option that was obviously bad — it's the mediocre one that never had an exit written, so it lived on "a bit more data" until the sunk cost defended it. A kill criterion set before you care costs a sentence; the same decision made after months of investment costs an argument, and usually loses it. The model can draft those exits in a sentence, and honoring one when it fires stays with the people who own the option — but the cheapest time to decide when to quit is still before you've spent anything worth protecting.

Tools for this guide

Each generates the prompt described above — you run it in your own AI assistant.

Ready-made resources

Reusable prompts and templates for the exact steps in this guide.

Take it further

When this task is one step inside a larger workflow or build.

FAQ

Isn't setting kill criteria just being pessimistic about an option before it has a chance?

No — it's the opposite of giving up early. A kill criterion doesn't lower the bar or pre-decide failure; it names, honestly and in advance, what evidence would mean the option isn't working, so you can pursue it wholeheartedly knowing there's an agreed exit if the evidence turns. The alternative isn't optimism — it's an option with no failure condition, which is how a mediocre idea survives on "a bit more data" long after it stopped earning its place. Written well, kill criteria let you commit harder, not less, because you're not secretly wondering when to pull the plug: you already decided what would tell you.

Does NewPrompt track the metrics or kill the option when a criterion is hit?

No. NewPrompt helps you draft the kill criteria — the triggers, thresholds, deadlines, and owners — but it doesn't monitor a metric, watch a calendar, send a reminder, run an experiment, or retire an option when a threshold is crossed. The criteria go into your own tracker, dashboard, or review process, and checking them on the dates you set is a person's job. And when a trigger does fire, NewPrompt doesn't decide anything: the owner and stakeholders look at what fired and make the continue, kill, or defer call. The platform structures the thinking; the tracking and the decision live entirely on your side.

The AI proposed a threshold — should I just use the number it gave?

Treat it as a draft, not a verdict. The model can suggest a reasonable-looking threshold, but it doesn't know your real constraints — your actual runway, risk tolerance, or what "good enough" means for your users — so a number it invented can be too loose, too strict, or aimed at the wrong signal. The useful thing it does is surface the assumptions and give you a concrete starting point to argue with. Sharpen every threshold against your own context, confirm the evidence is something you can actually observe, and get the owner to agree the bar before the option proceeds — the criteria have to be yours to be honest.

How is this different from a decision matrix or deciding when to defer a decision?

By what they do and when. A decision matrix and a defer check both come before you commit — one picks the best option, the other decides whether it's even time to choose. Kill criteria come after: they assume you're already pursuing an option and ask what evidence, agreed up front, would make you stop. So they aren't competitors but a sequence — compare, pick, then set the criteria that would tell you the pick was wrong — and kill criteria own the part the other two don't touch: getting out of a choice past the point it stopped earning its place.