KI Product Management 15. September 2026 · 8 Min Lesezeit · Teil 2 von 2

Die Spezifikation ist der Rahmen. Was Produktteams aufschreiben müssen, damit KI allein arbeiten darf.

TL;DR

Kontext, Harness und Loop sagen, wo ein Problem sitzt. Womit man sie füllt, sagt die Spezifikation. Spec-Driven Development macht sie zur verbindlichen Quelle statt zur Vorstufe, und der wirksame Teil daran ist nicht die Beschreibung, sondern die Abnahmebedingung. Wer sie nicht formulieren kann, kann nicht delegieren, an keinen Agenten und an keinen Menschen. Die Grenze liegt dort, wo Verstehen erst beim Bauen entsteht: Für Erkundung gehört kein Kriterium auf den Zettel, sondern ein Lernziel.

Regel zum Start

Bevor du weiterliest: mach den Vier-Zeiler zur Regel.

Kopier diesen Prompt in die Dauereinstellungen deines KI-Werkzeugs, also Custom Instructions, Projektregeln oder CLAUDE.md. Er verlangt vor jeder Aufgabe vier ausgefüllte Zeilen und fragt nach, solange eine fehlt. Warum die vierte Zeile die anderen drei trägt, steht weiter unten.

Dauerregel für unsere Zusammenarbeit: Bevor du eine Aufgabe ausführst,
stehen vier Zeilen.

1. Ziel: ein Satz im Präsens, der den Zustand beschreibt statt der
   Tätigkeit. Bei Bewertungen gehört die Ergebnisform hinein, also
   Optionen mit Empfehlung statt fertiger Entscheidung.
2. Nicht dabei: zwei bis drei Dinge, die naheliegen und trotzdem
   draußen bleiben.
3. Grundlage: Quellen, Daten, Vorgaben. Was fehlt, benennst du als
   Lücke und füllst es nicht.
4. Fertig, wenn: eine Beobachtung, die das Gegenteil beweisen würde.
   Sonst ist es kein Kriterium.

Geh so vor:
1. Schreib die vier Zeilen aus meiner Aufgabe heraus, bevor du anfängst.
   Was ich nicht gesagt habe, bleibt leer. Rate es nicht.
2. Frag nach den leeren Zeilen, höchstens zwei Fragen auf einmal.
3. Prüf die vierte Zeile gegen zwei Tests: Gibt es eine Beobachtung, die
   sie verletzt? Und kann das jemand anderes feststellen als ich? Wenn
   einer der beiden Tests durchfällt, schlag eine schärfere Fassung vor.
4. Erst wenn die vier Zeilen stehen, fängst du an. Am Ende prüfst du dein
   Ergebnis gegen die vierte Zeile und sagst mir, woran es scheitert.

Ausnahme: Wenn ich noch nicht weiß, wie das Ergebnis aussehen muss, ist
es Erkundung. Dann tritt an die Stelle der vierten Zeile ein Lernziel und
du fragst mich, was ich danach entscheiden können will.

Der erste Teil hat drei Ebenen sortiert: Kontext legt fest, was ein Modell für einen einzelnen Aufruf sieht, das Harness die Umgebung eines Agenten, der Loop, wer ihn anstößt.1 Das hilft bei der Fehlersuche und beantwortet die drängendere Frage nicht. Wer eine KI-gestützte Funktion beschreiben soll, sitzt vor einem leeren Dokument. Die Ebenen sagen, wo etwas hingehört. Sie sagen nicht, was hineingehört.

Was Spec-Driven Development umdreht

Für diese Lücke gibt es einen Ansatz, der die übliche Reihenfolge umkehrt. Der Gedanke selbst ist alt. Anforderungsdokumente, Design Docs und Design by Contract stellen die Beschreibung seit langem vor die Umsetzung.2 Neu ist nicht der Gedanke, sondern sein Adressat. Die Spezifikation richtet sich nicht mehr an Menschen, die nachfragen können, sondern an Agenten, die es nicht tun.

Vor der Umsetzung steht eine Spezifikation und diese Spezifikation ist nicht der Entwurf der Wahrheit, sondern die Wahrheit selbst. GitHub beschreibt die Verschiebung als Wechsel von „Code ist die Quelle der Wahrheit" zu „Absicht ist die Quelle der Wahrheit" und schlägt einen Ablauf in vier Schritten vor: spezifizieren, planen, in Aufgaben zerlegen, umsetzen.3 Jeder Schritt erzeugt ein Dokument, das den nächsten speist.

Birgitta Böckeler beschreibt das Artefakt genauer: ein strukturiertes, verhaltensorientiertes Dokument in natürlicher Sprache, das Mensch und Agent gemeinsam als Bezugspunkt nutzen.4 Sie unterscheidet drei Ausbaustufen, die oft vermischt werden: Die Spezifikation stößt die Arbeit nur an, sie läuft synchron mit dem Code weiter oder sie ist das einzige Artefakt, das noch von Hand bearbeitet wird. Der Unterschied entscheidet, welches Dokument bei einem Widerspruch gewinnt.

Warum das eine Produktfrage ist

Kief Morris beschreibt drei Haltungen, die ein Team gegenüber einem Agenten einnehmen kann.5

1

Außerhalb der Schleife. Ein Ziel nennen und das Ergebnis nicht prüfen.

2

In der Schleife. Jedes Ergebnis einzeln abnehmen. Damit wird der Mensch selbst zum Engpass, weil Agenten schneller erzeugen als Menschen lesen.

3

An der Schleife. Nicht einzelne Ergebnisse abnehmen, sondern die Spezifikationen und Prüfungen bauen, an denen sich jedes Ergebnis messen lassen muss.

Nur die dritte Haltung skaliert. Sie verschiebt die Arbeit dorthin, wo Produktleute ohnehin arbeiten. Wer an der Schleife steht, schreibt keine Prompts, sondern Bedingungen. Das ist dieselbe Tätigkeit wie das Formulieren einer Definition of Done.

Zwei Dinge ändern trotzdem alles. Die Bedingung greift bei jeder einzelnen Ausführung statt einmal pro Story. Und sie steht vor der Arbeit statt im Kommentarfeld darunter.

Autonomie entsteht nicht dadurch, dass ein Agent mehr darf, sondern dadurch, dass vorher feststeht, woran sein Ergebnis scheitern kann.

Die vier Felder, die eine Spezifikation tragen

Daraus lässt sich ein knapper Satz an Feldern ableiten. Er ist bewusst kurz, weil Vollständigkeit hier kein Qualitätsmerkmal ist. Eine lange Spezifikation erhöht die Prüflast, ohne die Ergebnisse besser zu machen.

Feld Beantwortet Ohne dieses Feld passiert
Absicht Was soll gelten, und wogegen argumentiert das Ergebnis Der Agent liefert etwas Plausibles, das die eigentliche These verfehlt
Geltungsbereich Was gehört dazu und was ausdrücklich nicht Themenwanderung, der Umfang wächst während der Ausführung
Zulieferung Welche Quellen, Daten und Vorgaben verwendet werden dürfen Der Agent füllt Lücken mit Plausiblem statt mit Belegtem
Abnahmebedingung Woran das Ergebnis geprüft wird und woran es scheitern darf Autonomie ist nicht begrenzt, sondern nur unbeobachtet

Das vierte Feld trägt die anderen drei. Absicht, Geltungsbereich und Zulieferung wirken vor der Ausführung und lassen sich beliebig ausschmücken, ohne dass sich am Ergebnis etwas ändert. Die Abnahmebedingung wirkt danach und korrigiert als einzige den Agenten. Eine ausführliche Beschreibung ohne sie produziert zuverlässig gut aussehende Ergebnisse, die niemand geprüft hat.

Notizblatt Der Vier-Zeiler

Die vier Felder passen auf einen Zettel. In dieser Reihenfolge ausfüllen, bevor die Aufgabe rausgeht, an eine KI wie an einen Menschen.

Ziel

Ein Satz im Präsens, der den Zustand beschreibt statt der Tätigkeit. Bei Bewertungen gehört die Ergebnisform hinein: Optionen mit Empfehlung, nicht die fertige Entscheidung.

Nicht dabei

Zwei bis drei Dinge, die naheliegen und trotzdem draußen bleiben.

Grundlage

Quellen, Daten, Vorgaben. Was fehlt, wird als Lücke benannt und nicht gefüllt, sonst wird es erfunden.

Fertig, wenn

Eine Beobachtung, die das Gegenteil beweisen würde. Sonst ist es kein Kriterium.

Ausgefüllt, erstes Beispiel: ein Rechercheauftrag

Ziel Die Wettbewerbsübersicht zeigt für fünf Anbieter Preis, Zielgruppe und Alleinstellung.

Nicht dabei Bewertung, Empfehlung, Priorisierung.

Grundlage Nur die öffentlichen Preisseiten der fünf Anbieter, mit Abrufdatum.

Fertig, wenn Jede Zahl trägt einen Link auf die Seite, auf der sie steht. Der Link führt tatsächlich dorthin.

Ausgefüllt, zweites Beispiel: ein Gesprächsleitfaden

Ziel Der Leitfaden bringt eine Person dazu, von ihrem letzten echten Fall zu erzählen, in ihrer eigenen Reihenfolge.

Nicht dabei Fragen nach Lösungen, Fragen nach Wünschen und jede Formulierung, die eine meiner Hypothesen schon enthält.

Grundlage Die Rolle der Person und die zwei Annahmen, die das Gespräch kippen könnten. Fehlt die Rolle, wird der Leitfaden nicht gebaut.

Fertig, wenn Keine Frage lässt sich mit ja oder nein beantworten, und keine nennt einen Begriff, den die Person nicht selbst benutzt hat.

Bei diesem Beispiel tut die zweite Zeile weh. Genau deshalb steht es hier. Die eigenen Ideen draußen zu lassen ist die schwerste Zeile, die ein Produktmensch schreiben kann. Ich habe das an einem Dienstag im September gelernt, an dem zwei fertige Leitfäden in den Papierkorb gingen, weil ich die Rolle des Gegenübers nicht mitgeliefert hatte.

Die vierte Zeile ist die einzige, die Arbeit spart. Wer sie nicht hinbekommt, hat noch keine Aufgabe, sondern eine Stimmung.

Und die vierte Zeile hat eine Verschärfung, die leicht zu übersehen ist. Eine Abnahmebedingung, die nur ein Mensch prüfen kann, verschiebt die Prüflast. Sie senkt sie nicht. Die Frage hinter der vierten Zeile lautet deshalb nicht nur „woran scheitert das Ergebnis", sondern „wer stellt das fest, ohne dass ich hinsehe".

Ob jede Zahl einen Link trägt und der Link dorthin führt, prüft ein kleines Skript, bevor ich den Text öffne. Ob ein Text gut belegt ist, prüfe nur ich, jedes Mal neu. Zwei Bedingungen, die dasselbe meinen. Nur eine macht mich frei.

Kritische Einordnung

Der schwerwiegendste Einwand kommt aus der Nachbarschaft. Martin Fowler gibt Kent Becks Kritik wieder: Die Annahme, Anforderungen ließen sich vor der Umsetzung vollständig festlegen, ignoriert, dass beim Bauen Erkenntnisse entstehen, die die Spezifikation verändern würden.6 Wer sie zur alleinigen Wahrheit erklärt, definiert genau dieses Lernen weg. Der Einwand trifft die dritte Ausbaustufe härter als die ersten beiden.

Wann keine Abnahmebedingung

Becks Einwand lässt sich auflösen, aber nicht durch eine bessere Spezifikation. Er lässt sich auflösen, indem man aufhört, alles gleich zu behandeln.

Für wiederkehrende Aufgaben mit bekanntem Ergebnis schreibt man Abnahmebedingungen. Für Erkundung schreibt man Lernziele. Das ist kein Kompromiss zwischen beidem, es sind zwei verschiedene Dokumente für zwei verschiedene Lagen.

Wer eine Discovery mit einer Abnahmebedingung belegt, hat sie zur Lieferung erklärt und bekommt genau das zurück: ein Ergebnis, das die Bedingung erfüllt, und keine einzige Überraschung. Bei einer Wettbewerbsübersicht ist das richtig, bei fünf Nutzergesprächen ein Eigentor. Die Bedingung „drei Muster identifiziert" erzeugt drei Muster, auch wenn keine da sind.

Ein Lernziel sieht anders aus. Es nennt die Annahme, die geprüft wird, und den Satz, der sie kippen würde. Es nennt kein Ergebnis, weil das Ergebnis der Punkt der Übung ist. Geprüft wird nicht, ob geliefert wurde, sondern ob ich hinterher mehr weiß.

Die Unterscheidung ist leicht und wird trotzdem selten getroffen: Weiß ich, wie das Ergebnis aussehen muss? Dann Abnahmebedingung. Will ich es herausfinden? Dann Lernziel. Im Zweifel Erkundung, denn eine überflüssige Bedingung kostet Zeit, eine falsche die Erkenntnis.

Dazu kommen zwei praktische Beobachtungen von Böckeler.4 Die Menge an Spezifikationsdokumenten kann die Prüflast erhöhen statt sie zu senken, weil am Ende jemand alle lesen muss. Und Agenten ignorieren Vorgaben auch dann, wenn diese ausführlich niedergeschrieben sind.

Wie weit sich das treiben lässt, zeigt BMAD-METHOD. Die Werkzeugsammlung verteilt die Arbeit auf benannte Agenten-Personas, darunter einen Product Manager namens John, dessen Arbeitsablauf „create-prd“ heißt.7 Die Produktrolle ist dort als Prompt implementiert. Damit lautet die Frage nicht mehr, wie man eine Spezifikation schreibt, sondern wer für sie geradesteht.

Wie das Feld den Ansatz einordnet

Der schärfste Einwand trifft nicht die Substanz, sondern den Neuheitsanspruch. Brandon Kindred führt den Ansatz auf den Teil des Entwicklungszyklus zurück, den die meisten Organisationen seit Jahren überspringen. Sein Fazit: KI habe Spec-Driven Development nicht erfunden, sondern das Überspringen teuer gemacht.2

Roger Wong8 und Yuval Yeret9 antworten beide auf den Wasserfall-Vorwurf statt auf die Frage, ob der Ansatz trägt.

Der nüchternste Befund kommt aus dem Haus der hier zweimal zustimmend zitierten Quelle. Böckeler arbeitet bei Thoughtworks. Thoughtworks führt den Ansatz im eigenen Technology Radar unter „Assess“ statt unter „Adopt“.10 Das ist die Einstufung für Ansehen, nicht für Einsetzen.

Den Rest sagt das Tempo, mit dem der Ansatz durch die Disziplinen wandert. Das Design beansprucht ihn bereits für sich,11 die Wissenschaft für reproduzierbare Datenanalyse.12 Wer ihn für die Produktarbeit reklamiert, reiht sich ein, statt etwas zu eröffnen.

Der Übertrag in den Produktalltag

Der Ansatz ist nicht auf Softwareentwicklung beschränkt. Für Produktarbeit heißt er: Bevor eine wiederkehrende Aufgabe an eine KI geht, wird sie einmal spezifiziert, mit einer Bedingung, an der das Ergebnis scheitern darf.

Der Aufwand fällt einmal an, die Prüfung dauerhaft weg. Eine Spezifikation, die einmal geschrieben und danach nachgeschärft wird, ersetzt die Rückfragen aus jedem Durchlauf. Der Bruch liegt dort, wo sie nur im Kopf existiert oder in einem Dokument, das niemand findet.

Die Bedingung gehört vor die Arbeit, nicht dahinter. Eine Abnahmebedingung, die erst formuliert wird, wenn das Ergebnis vorliegt, ist keine Bedingung, sondern eine nachträgliche Begründung.

Was bleibt für jeden Tag

Erst die Abnahmebedingung, dann die Beschreibung

Wer nicht sagen kann, woran ein Ergebnis scheitern darf, hat keine Aufgabe formuliert, sondern einen Wunsch.

Prüfbar heißt: es gibt eine Beobachtung, die sie verletzt

Der Test funktioniert auch für menschliche Abnahmen und deckt schwammige Kriterien in Sekunden auf.

Und dann die Gegenprobe: ist das hier überhaupt eine Lieferung?

Wenn ich noch nicht weiß, wie das Ergebnis aussehen muss, brauche ich kein Kriterium, sondern ein Lernziel.

Die Spezifikation ist kein Ersatz für das Prüfen

Sie macht das Prüfen möglich und billig. Wer sie schreibt und dann nicht hinsieht, hat den Aufwand umsonst betrieben.

Quellen

  1. Scharffetter, L. (2026). Kontext, Harness, Loop. Drei Ebenen, die Produktteams gerade durcheinanderbringen. linja.me (Teil 1 dieses Zweiteilers).
  2. Kindred, B. (2026). Same Patterns, New Hype: Spec-Driven Development. Medium, 20. April 2026.
  3. Delimarsky, D. (2025). Spec-driven development with AI: Get started with a new open source toolkit. The GitHub Blog, 2. September 2025.
  4. Böckeler, B. (2025). Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl. martinfowler.com, 15. Oktober 2025.
  5. Morris, K. (2026). Humans and Agents in Software Engineering Loops. martinfowler.com, 4. März 2026.
  6. Fowler, M. (2026). Fragments: January 8. martinfowler.com, 8. Januar 2026. (Gibt Kent Becks Einwand gegen vorab vollständige Spezifikationen wieder.)
  7. BMad Method (o. J.). Agents. Projektdokumentation, bmad-code-org/BMAD-METHOD. Abgerufen am 16. September 2026.
  8. Wong, R. (2026). Spec-Driven Development: It Looks Like Waterfall (And I Feel Fine). rogerwong.me, 4. März 2026.
  9. Yeret, Y. (2026). Is Spec-Driven Development a Step Forward or Back? yuvalyeret.com, 5. Juni 2026.
  10. Thoughtworks (2025). Spec-driven development. Technology Radar Vol. 34, November 2025, Ring „Assess“.
  11. Härkönen, T. (2026). Faster prototypes miss the point: The case for spec-driven design. futurice.com, 11. August 2026.
  12. Chen, C., Luo, B., Li, N., Wang, B., Yang, H., Guo, J. & Xu, M. (2025). Spec-Driven AI for Science: The ARIA Framework for Automated and Reproducible Data Analysis. arXiv:2510.11143, 13. Oktober 2025.

Glossar

Spec-Driven Development — Arbeitsweise, bei der eine Spezifikation vor der Umsetzung steht und als verbindliche Quelle gilt.

Spezifikation — strukturiertes Dokument in natürlicher Sprache, das beschreibt, was gelten soll, nicht wie es hergestellt wird.

Abnahmebedingung — vorab formuliertes, prüfbares Kriterium, an dem ein Ergebnis scheitern kann.

Definition of Done — im agilen Arbeiten die verbindliche Liste, wann eine Aufgabe als fertig gilt.

Agent — ein Sprachmodell, das über mehrere Schritte hinweg Werkzeuge benutzt und Zustand behält.

Harness — alles an einem KI-Agenten, was nicht das Modell ist: Werkzeuge, Rechte, Zustand, Prüfungen.

Loop — das System, das einen Agenten wiederholt anstößt, prüft und weiterlaufen lässt.

Kontext — die Informationen, die ein Modell für einen einzelnen Aufruf sieht.

Agentische KI — KI-Systeme, die eigenständig Handlungsschritte ausführen statt nur zu antworten.

Prüflast — der Aufwand, den Menschen aufbringen müssen, um Ergebnisse eines Systems zu kontrollieren.

Lernziel — Gegenstück zur Abnahmebedingung: benennt die Annahme, die geprüft wird, und den Befund, der sie kippen würde, statt ein Ergebnis vorzugeben.

Discovery — die Phase der Produktarbeit, in der ein Problem noch untersucht und nicht bereits umgesetzt wird.

Design by Contract — Entwurfsprinzip, bei dem ein Bauteil vorab zusichert, was es erwartet und was es garantiert.

RFC — Request for Comments: Vorschlagsdokument, das eine geplante Änderung beschreibt und zur Kommentierung stellt.

PRD — Product Requirements Document: das klassische Dokument, in dem Produktanforderungen festgehalten werden.

Technology Radar — halbjährliche Einordnung von Technologien durch Thoughtworks in die Ringe Adopt, Trial, Assess und Hold.

Teil 1

Die Ebenen, auf die sich dieser Teil bezieht, stehen in Kontext, Harness, Loop: was die drei Begriffe trennt, wie sie aufeinander aufbauen und woran man erkennt, auf welcher Ebene ein Problem sitzt.

AI Product Management September 15, 2026 · 8 min read · Part 2 of 2

The specification is the frame. What product teams have to write down before AI is allowed to work alone.

TL;DR

Context, harness and loop say where a problem sits. What fills them is the specification. Spec-driven development makes it binding rather than a draft, and the part that works is not the description but the acceptance condition. Anyone who cannot state one cannot delegate, to an agent or a person. The limit sits where understanding only emerges while building: exploration needs a learning goal, not a criterion.

Rule first

Before you read on: turn the four-liner into a rule.

Paste this prompt into the standing settings of your AI tool, meaning custom instructions, project rules or CLAUDE.md. It requires four filled lines before any task and keeps asking while one is missing. Why the fourth line carries the other three is further down.

Standing rule for how we work: before you carry out a task, four lines
are in place.

1. Goal: one sentence in the present tense describing the state rather
   than the activity. For assessments the output form belongs in it,
   meaning options with a recommendation instead of a finished decision.
2. Not included: two or three things that suggest themselves and stay
   out anyway.
3. Basis: sources, data, constraints. Whatever is missing you name as a
   gap instead of filling it.
4. Done when: an observation that would prove the opposite. Without one
   it is not a criterion.

Work like this:
1. Write the four lines out of my task before you start. Whatever I did
   not say stays empty. Do not guess it.
2. Ask me about the empty lines, at most two questions at a time.
3. Put the fourth line through two tests: is there an observation that
   violates it? And can anyone other than me establish that? If either
   test fails, propose a sharper version.
4. Only start once the four lines stand. At the end, check your result
   against the fourth line and tell me where it fails.

Exception: if I do not yet know what the result has to look like, this is
exploration. Then a learning goal takes the place of the fourth line and
you ask me what I want to be able to decide afterwards.

Part one sorted three layers: context settles what a model sees for a single call, the harness the environment an agent acts in, the loop who triggers it.1 That helps with debugging and leaves the more pressing question open. Anyone asked to describe an AI supported feature sits in front of an empty document. The layers say where something belongs. They do not say what goes in.

What spec-driven development reverses

There is an approach that inverts the usual order, and the idea itself is old. Requirements documents, design docs and design by contract have long put the description before the build.2 What is new is not the idea but its addressee. The specification no longer speaks to people who can ask, but to agents that will not.

The specification comes before the build, and it is not a draft of the truth but the truth itself. GitHub frames the shift as moving from "code is the source of truth" to "intent is the source of truth", in four steps: specify, plan, break into tasks, implement.3 Each step produces a document feeding the next.

Böckeler describes the artefact more precisely: a structured, behaviour oriented document in natural language that human and agent share as a reference point.4 She separates three levels of ambition the debate tends to blur: the specification only kicks the work off, it evolves in sync with the code, or it is the only artefact edited by hand. The difference settles which document wins when they contradict.

Why this is a product question

Kief Morris describes three stances a team can take towards an agent.5

1

Outside the loop. Name a goal and never check the result.

2

In the loop. Sign off every output individually. That makes the human the bottleneck, because agents generate faster than humans read.

3

On the loop. Do not sign off individual outputs. Build the specifications and checks every output has to survive.

Only the third stance scales: it moves the work to where product people already are. Standing on the loop means writing conditions, not prompts. That is the same activity as writing a definition of done.

Two things change everything anyway. The condition applies to every run instead of once per story. And it sits before the work, not in the comment field below it.

Autonomy does not come from allowing an agent more. It comes from settling beforehand what its result is allowed to fail against.

The four fields that carry a specification

From those sources and the layer logic of part one, a short set of fields follows, deliberately short because completeness is not a quality marker here. A long specification raises the review load without improving the output.

Field Answers Without it you get
Intent What should hold, and what the result argues against Something plausible that misses the actual claim
Scope What belongs in and what explicitly does not Topic drift, scope growing during the run
Inputs Which sources, data and constraints may be used Gaps filled with the plausible instead of the evidenced
Acceptance condition What the result is checked against and may fail against Autonomy that is not bounded, merely unobserved

The fourth field carries the other three. Intent, scope and inputs act before the run and can be embellished endlessly without changing the output. The acceptance condition acts afterwards and is the only one correcting the agent. A rich description without it reliably produces good looking results nobody has checked.

Notepad The four-liner

The four fields fit on one sheet. Fill them in, in this order, before the task goes out, to an AI as much as to a person.

Goal

One sentence in the present tense, describing a state rather than an activity. For anything evaluative, name the form of the answer: options with a recommendation, not the decision itself.

Not included

Two or three things that look obvious and stay out anyway.

Basis

Sources, data, constraints. Whatever is missing gets named as a gap and left open, otherwise it gets invented.

Done when

An observation that would prove the opposite. Otherwise it is not a criterion.

Filled in, first example: a research assignment

Goal The competitor overview shows price, target group and differentiator for five vendors.

Not included Assessment, recommendation, prioritisation.

Basis Only the public pricing pages of those five vendors, with the date they were retrieved.

Done when Every number carries a link to the page it sits on. That link actually leads there.

Filled in, second example: an interview guide

Goal The guide gets a person to tell the story of their last real case, in their own order.

Not included Questions about solutions, questions about wishes, and any wording that already carries one of my hypotheses.

Basis The person's role and the two assumptions the conversation could overturn. Without the role, the guide does not get built.

Done when No question can be answered with yes or no, and none uses a term the person has not used themselves.

In this example the second line hurts, which is why it is here. Leaving your own ideas out is the hardest line a product person can write. I learned that on a Tuesday in September, when two finished interview guides went in the bin because I had not supplied the role of the person opposite.

The fourth line is the only one that saves work. Anyone who cannot write it does not yet have a task, only a mood.

And the fourth line carries a sharpening that is easy to miss. An acceptance condition only a human can check moves the review load, it does not lower it. The question is not just what the result fails against, but who establishes that without me looking.

A script checks whether every number's link leads where it claims, before I open the text. Whether a text is well evidenced, only I can check, every time. Two conditions, one meaning. Only one frees me.

A critical look

The heaviest objection comes from close by. Martin Fowler relays Kent Beck's criticism: assuming requirements can be settled fully before implementation ignores that building produces insights that would change the specification.6 Declaring it the sole truth risks defining that learning away. The objection hits the third level of ambition hardest.

When not to write an acceptance condition

Beck's objection can be resolved, but not by a better specification. It resolves once you stop treating every task the same way.

For recurring tasks with a known result, write acceptance conditions. For exploration, write learning goals. Not a compromise between the two, but two documents for two situations.

Putting an acceptance condition on a discovery declares it a delivery, and that is what comes back: a result that meets the condition, and no surprises. Right for a competitor overview, an own goal for five user interviews. The condition "three patterns identified" produces three patterns, even when there are none.

A learning goal looks different. It names the assumption under test and the sentence that would overturn it, but no result, because the result is the point. The check is not whether something was delivered, but whether I know more afterwards.

The distinction is easy and still rarely made. If you know what the result has to look like, write an acceptance condition. If you want to find out, write a learning goal. In doubt, treat it as exploration: a redundant condition costs time, a wrong one the insight.

Two practical observations from Böckeler point the same way.4 The volume of specification documents can raise the review load instead of lowering it, because somebody has to read all of them. And agents ignore instructions even when those are written out in detail.

How far that can be taken shows in BMAD-METHOD. The toolkit splits the work across named agent personas, among them a product manager called John, whose workflow is named “create-prd”.7 The product role is implemented there as a prompt. The question stops being how to write a specification and becomes who answers for it.

How the field rates the approach

The sharpest objection misses the substance and hits the claim to novelty. Brandon Kindred traces the approach back to the part of the development lifecycle most organisations have been skipping for years. His verdict: AI did not invent spec-driven development but made skipping it expensive.2

Roger Wong8 and Yuval Yeret9 both answer the waterfall charge rather than the question of whether the approach holds.

The soberest reading comes from the house of the source quoted approvingly twice above. Böckeler works at Thoughtworks. Thoughtworks lists the approach in its own Technology Radar under “Assess” rather than “Adopt”.10 That is the rating for looking, not for using.

The rest is said by how fast the approach travels across disciplines. Design already claims it,11 science claims it for reproducible data analysis.12 Claiming it for product work means joining a queue, not opening a door.

What this changes in product work

The approach is not limited to software engineering. For product work it reads: before a recurring task goes to an AI, it gets specified once, with a condition the result is allowed to fail against.

The effort lands once, the questions go away permanently. A specification written once and sharpened afterwards replaces the follow-ups from every run. It breaks down where it exists only in someone's head or in a document nobody finds.

The condition belongs before the work, not after it. An acceptance condition written once the result is in is not a condition but a retrospective justification.

What stays with you every day

Acceptance condition first, description second

If you cannot say what a result is allowed to fail against, you have not stated a task but a wish.

Checkable means there is an observation that violates it

The test works for human sign-off too and exposes vague criteria in seconds.

Then the counter-check: is this a delivery at all

If you do not yet know what the result has to look like, you need a learning goal, not a criterion.

A specification is not a substitute for checking

It makes checking possible and cheap. Writing one and not looking wastes the effort.

Glossary

Spec-driven development — working from a specification written before the build and treated as the binding source.

Specification — a structured natural language document stating what should hold, not how it gets built.

Acceptance condition — a checkable criterion stated up front that a result is allowed to fail against.

Definition of done — in agile work, the binding list of what makes a task finished.

Agent — a language model that uses tools across several steps and keeps state.

Harness — everything about an AI agent that is not the model: tools, permissions, state, checks.

Loop — the system that triggers an agent repeatedly, verifies it and keeps it running.

Context — the information a model sees for a single call.

Agentic AI — AI systems that carry out steps on their own instead of only answering.

Review load — the effort humans have to spend checking what a system produced.

Learning goal — the counterpart to an acceptance condition: it names the assumption under test and the finding that would overturn it, instead of prescribing a result.

Discovery — the phase of product work in which a problem is still being investigated rather than built.

Design by contract — a design principle where a component states up front what it expects and what it guarantees.

RFC — request for comments: a proposal document describing a planned change and opening it up for comment.

PRD — product requirements document: the classic artefact recording what a product needs to do.

Technology Radar — Thoughtworks' twice yearly placement of technologies into the rings adopt, trial, assess and hold.

Part 1

The layers this part builds on are in Context, harness, loop: what separates the three terms, how they stack, and how to tell which layer a problem sits on.