Features/MCP Workflows/Google Docs/Customer specification extraction

Customer specification extraction

Half of any customer specification is scope, definitions and narrative, and none of it may become a requirement. The hard part is drawing that line, so this shows you the line before it imports anything.

Google DocsWrites · gated

The problem this solves

Customer specifications arrive as prose where nothing is labelled. Extract too eagerly and the model fills with background statements that nobody can verify; extract too cautiously and a binding obligation is missed. Either way the wording gets tidied on the way in, and the model no longer says what the contract says.

How it works

Requirements out of a customer spec, with what you deliberately left out listed beside them. Half of any customer specification is scope, definitions and narrative. The extraction rule is approved against fifteen real examples before anything is imported, and every requirement carries its section, page and document revision.

01
Agree the extraction rule on real examples
Which verbal forms bind, how the identifier scheme reads, which sections carry requirements. Shown with ten real statements it would extract and five it would not, each with the reason, before anything is imported.
02
Preserve the contractual wording
Identifiers exactly as written and statements verbatim: no paraphrase, no grammar fixes, no expanded abbreviations, no sentences split for readability. Document revision, section, page and surrounding context recorded on each.
03
Flag rather than resolve
Compound paragraphs, unnumbered statements, TBDs, near-duplicates, pointers to other documents and internal conflicts. Splitting a compound requirement changes the contractual unit, so that is your decision.
04
List what it left out
Scope, definitions, background, references and informative annexes, grouped by reason, so the boundary it drew can be audited rather than taken on trust.

The prompt

You are extracting requirements from a customer specification in
Google Docs into the Dalus model [model name]. The hard part is not
copying text; it is deciding what is a requirement and what is
context, in a document where nothing is labelled. Roughly half of any
customer specification is scope, background, definitions and
narrative, and none of that may become a requirement. Everything you
extract must be traceable back to where it came from, because this
material is contractual. This workflow writes to the model, and only
after an approved dry run.

CONFIRM FIRST, in one message: which Dalus model; which Google Doc,
and its revision or date; whether the whole document or named
sections are in scope; where the extracted requirements should sit in
the model; and whether this is a first extraction or a re-extraction
after the customer issued a new revision. Ask anything else in the
same message. Then begin.

READ THE DOCUMENT COMPLETELY before extracting anything: its section
structure, numbering scheme, any identifier convention it uses, its
verbal forms, and any table of requirements it contains. Note the
document revision, date and title for provenance.

STATE YOUR EXTRACTION RULE AND GET IT APPROVED before importing.
Show: which verbal forms you are treating as binding (shall, must,
will, is to be, and whatever the document actually uses); how you
are reading its identifier scheme; which sections you will treat as
requirement-bearing and which as context; and how you are handling
tables, figures and notes. Then show fifteen real examples from the
document: ten you would extract and five you would not, each with
the reason. A wrong rule applied to a 200-page specification produces
a mess that is expensive to unpick, and the rule is the user's
decision.

WHAT NOT TO EXTRACT, reported explicitly rather than silently
dropped: scope and purpose statements, definitions and abbreviations,
background and rationale, descriptions of the customer's existing
system, references to other documents, and anything in a note or
informative annex that the document itself marks as non-binding.
Deliver this as a list of what you deliberately left out, grouped by
reason, so the boundary you drew can be audited. Where the document
uses shall-language inside a section that is clearly context, flag it
rather than deciding alone.

PROVENANCE ON EVERY EXTRACTED REQUIREMENT, without exception:
- The original identifier, preserved exactly, as the customer ID.
  Never renumber, never normalise the format.
- The statement, verbatim. Never paraphrase, never correct grammar,
  never expand abbreviations, never split a sentence for
  readability. This is contractual wording.
- The document title, revision and date.
- The section number and heading it came from, and the page.
- The surrounding text, enough to read it in context later.
Set the requirement type from the section it came from, and say what
mapping you used from section to type.

FLAG, DO NOT SILENTLY RESOLVE:
- Compound requirements: one numbered paragraph containing several
  binding statements. Splitting changes the contractual unit, so
  report each one with the count of statements found and let the
  user decide between importing whole or splitting.
- Unnumbered requirements: binding statements with no identifier.
  Propose an identifier scheme and get it approved, keeping them
  visibly distinct from customer-numbered ones.
- Requirements with TBD, TBC or blank values.
- Duplicate or near-duplicate statements at different identifiers.
- Statements that reference another document for their actual
  content, which are pointers rather than requirements.
- Conflicting requirements within the document.

DRY RUN, before any write: the extraction rule as approved; counts
by section and by type; the first five and last five extracted
requirements in full with their provenance; every flagged item; the
deliberately-excluded list; and any collision with customer IDs
already in the model. Reconciliation arithmetic: binding statements
found = extracted + flagged for a decision + excluded, with reasons.
Suggest importing into a branch. Wait for approval.

ON RE-EXTRACTION after a new customer revision: match on the
customer ID, and report new, changed (with both texts), unchanged,
and requirements present in the model whose identifier no longer
appears in the new revision. Never delete a requirement because the
customer removed it; report it, because a removed requirement is a
scope change someone has to notice.

IMPORT the approved set. If anything fails partway, stop, report
what was created and what was not, and do not improvise.

FINAL REPORT: the arithmetic against actuals, the flagged items and
their decisions, the excluded list, and any quality findings about
the customer's document (ambiguity, untestable wording, conflicts)
as findings only, never as edits to the extracted text. Offer, do
not execute: proposing allocations for the extracted requirements,
and generating a compliance matrix against them.

Replace the [bracketed] placeholders with your model and project names.

What you get

Requirements with identifiers and wording preserved exactly
Provenance on each: document revision, section, heading, page, context
The flagged list: compound, unnumbered, TBD, duplicate, conflicting
The deliberately-excluded list, grouped by reason
Dalus + Google Docs

Reads the document and its revision, read-only. On a re-extraction after the customer issues a new revision, it matches on the customer ID and reports new, changed with both texts, unchanged, and identifiers that have disappeared. A requirement the customer removed is reported as a scope change, never deleted from the model.

Nothing is written until you approve it. The prompt carries the gate: the agent shows the full change set and waits, whether the target is the model or a system it reaches through a connector.

Common questions

How does it know what is a requirement?
It does not decide alone. It states its extraction rule and shows fifteen real statements from your document, ten it would take and five it would not, and waits for you to confirm. A wrong rule applied to a 200-page specification is expensive to unpick.
Will it split compound requirements?
Only if you tell it to. Splitting one numbered paragraph into three changes the contractual unit, so it reports the paragraph with the count of binding statements it found and lets you decide between importing whole or splitting.
What happens on the next customer revision?
It matches on the customer ID and reports what is new, what changed with both texts shown, what is unchanged, and which identifiers no longer appear. Nothing is deleted from the model, because a removed requirement is a scope change someone has to notice.