Features/MCP Workflows/SharePoint/SharePoint specification extraction

SharePoint specification extraction

Requirements extracted from a named version in a named library are worth more than requirements from a PDF someone emailed. This keeps that trace.

SharePointWrites · gated

The problem this solves

The specification is in a controlled library with a version history, and the requirements get retyped from a copy on somebody's desktop. The provenance is gone, and when the customer issues a new revision nobody can say which version the model was built from.

How it works

Requirements out of a controlled library, carrying the site, version and page they came from. Requirements from a named version in a named library are worth more than requirements from a PDF someone emailed. Where the library version and the revision printed inside the document disagree, that is reported as the control problem it is.

01
Capture the source before the content
Site, library, file, version, modified date and who last modified it, plus the document's own stated revision. Where SharePoint's version and the printed revision disagree, both are reported: that is a control problem worth knowing.
02
Agree the extraction rule on real examples
Ten statements it would extract and five it would not, each with the reason, before anything is imported.
03
Provenance on every requirement
Identifier exactly, statement verbatim, plus the site, library, file, version, section, heading, page and enough surrounding text to read it in context later.
04
Check the library for what is missing
The applicable documents the specification references, and which of them are actually present in the same site. Three missing out of five is a finding before anyone starts work.

What you need

Microsoft's ODSP remote server, or the community Softeria server in read-only mode with SharePoint presets. The Softeria route also covers Teams, so one install serves both.

The prompt

You are extracting requirements from a specification document stored
in SharePoint into the Dalus model [model name], preserving the trace
back to the controlled source. The reason this workflow exists rather
than a plain upload is that SharePoint holds the version history and
the document's controlled location, and requirements extracted from
"a PDF someone emailed me" are worth less than requirements extracted
from a named version in a named library.

CONFIRM FIRST, in one message: which Dalus model; which SharePoint
site and document library; which document, and which version if the
library holds several; whether the whole document or named sections
are in scope; where the requirements should sit in the model; and
whether this is a first extraction or a re-extraction after a new
version. Ask anything else in the same message. Then begin.

THIS WORKFLOW IS READ-ONLY ON SHAREPOINT. Read documents and their
metadata; never upload, edit, check out, or change permissions on
anything. A document library in an enterprise tenant is a controlled
store and often an audited one.

CAPTURE THE SOURCE PROPERLY, before reading the content: the site,
the library, the file name, the version number, the modified date and
who last modified it, and the document's own stated revision if it
carries one. Where SharePoint's version differs from the revision
printed inside the document, report both and flag the mismatch,
because that is a control problem the team will want to know about.

READ THE DOCUMENT COMPLETELY before extracting anything: its section
structure, numbering, any identifier convention, its verbal forms,
and any requirement tables. If the file is a scan or its tables do
not extract cleanly, name the sections you could not read reliably
and treat nothing from them as extracted.

STATE YOUR EXTRACTION RULE AND GET IT APPROVED. Show which verbal
forms you treat as binding, how you read the identifier scheme, which
sections are requirement-bearing and which are context, and how you
handle tables, figures and notes. Then show fifteen real examples:
ten you would extract and five you would not, each with the reason. A
wrong rule applied to a long specification produces a mess that is
expensive to unpick, and the rule is the user's decision.

WHAT NOT TO EXTRACT, reported rather than silently dropped: scope and
purpose statements, definitions, background and rationale,
descriptions of the customer's existing system, references to other
documents, and anything an informative annex marks as non-binding.
Deliver this as a list grouped by reason so the boundary you drew can
be audited. Where shall-language appears inside an obviously
contextual section, flag it rather than deciding alone.

PROVENANCE ON EVERY EXTRACTED REQUIREMENT, without exception: the
original identifier preserved exactly as the customer ID, never
renumbered; the statement verbatim, never paraphrased, corrected or
split; the SharePoint site, library, file and version; the section
number and heading; the page; and enough surrounding text to read it
in context later.

FLAG, DO NOT SILENTLY RESOLVE: compound requirements holding several
binding statements, where splitting changes the contractual unit and
is the user's call; unnumbered requirements, where an identifier
scheme must be proposed and approved; TBDs and TBCs; near-duplicates
at different identifiers; statements that point to another document
for their content; and conflicts within the document.

CHECK THE LIBRARY FOR RELATED DOCUMENTS the specification references,
and report which of them are present in the same site and which are
not. A specification that cites five applicable documents of which
three are missing from the library is a finding worth surfacing
before anyone starts work against it.

DRY RUN before any write: the source capture, the approved rule,
counts by section and type, the first five and last five extracted
requirements in full with provenance, every flagged item, the
excluded list, and any collision with customer IDs already in the
model. Reconciliation: binding statements found = extracted + flagged
+ excluded, with reasons. Suggest importing into a branch. Wait for
approval.

ON RE-EXTRACTION after a new document version: match on the customer
ID and report new, changed with both texts, unchanged, and
requirements in the model whose identifier no longer appears. Never
delete because the customer removed it; report it, since a removed
requirement is a scope change and may be a contractual one. Record
which SharePoint version each change came from.

FINAL REPORT: the arithmetic against actuals, the flagged items, the
excluded list, the missing referenced documents, and any quality
findings about the customer's document as findings only, never edits.
Offer, do not execute: proposing allocations for the extracted set,
and generating a compliance matrix against it.

Replace the [bracketed] placeholders with your model and project names.

What you get

Requirements with identifiers and wording preserved exactly
SharePoint site, library, version and page recorded per requirement
Version-versus-revision mismatches, flagged as control problems
Referenced documents absent from the library, listed
Dalus + SharePoint

Read-only on the tenant: it reads documents and their metadata and never uploads, edits, checks out or changes permissions on anything, since a document library in an enterprise tenant is a controlled and often audited store. On a re-extraction it records which SharePoint version each change came from.

Nothing is written until you approve it. The prompt carries the gate: the agent shows the full change set and waits, whether the target is the model or a system it reaches through a connector.

Common questions

Why not just upload the document?
Because SharePoint holds the version history and the controlled location, and that provenance is most of what makes an extracted requirement defensible later. Every requirement carries the site, library, file, version and page it came from.
What if the library version and the document's own revision disagree?
Both are reported and the mismatch is flagged. It is a control problem rather than an extraction problem, and the team will want to know before anyone works from either number.
Does it touch our library?
No. Read-only throughout: no uploads, no edits, no check-outs, no permission changes.