Archive · Tranche 2

Feeding the whole corpus to a machine

Fifty-four documents loaded into a research tool that generated briefing papers, study guides, data tables, slide decks and three AI-hosted audio discussions about a community pub campaign. Useful for onboarding, dangerous as a source.

Original title
NotebookLM — Bottom Pub Co-op Full Docs
Original date
May–June 2026 (last refreshed 12 June 2026)
Project phase
Whole campaign
Purpose at the time
Make a large and growing document corpus navigable for volunteers who could not read all of it, and generate meeting material without writing it by hand.
Status at the time
Live workspace, refreshed as documents changed. Twenty generated artefacts existed when it was last updated.
Source provenance
Derived from internal/notebooklm.md and the twenty generated artefacts under internal/notebooklm/artifacts/ (private working corpus; unpublished, unchanged).
Publication treatment
Derived
Derived / prepared by
Claude (Fable 5) with Adrian Wedd, 18 August 2026
Prepared
2026-08-18
Human review
Adrian Wedd — publication review completed 19 August 2026
Published
2026-08-19
What was changed for publication
  • Authored description. The workspace identifier, storage bucket hostname and the tooling configuration are withheld under carve-out 1.
  • The twenty generated artefacts are withheld as duplicative rather than sensitive: five briefing reports, five data tables, five slide decks, three audio discussions and two flashcard sets, all generated from documents that are themselves published in this archive. Publishing a machine's summary of a document alongside the document is noise.
  • The audio and slide material is roughly 150 MB and is the bulk of the corpus by size, which is worth knowing when reading the corpus map's file counts.

Document begins

Feeding the whole corpus to a machine

By June the project had more documents than any volunteer was going to read. Someone joining the finance stream needed the assumptions book, the feasibility draft, the capital scaffold and three case-study registers, and they had a job and a family.

So the corpus — 54 documents — was loaded into a research tool that reads a source set and generates material from it, and 20 artefacts were produced from it.

What came out

Artefact typeCountWhat it was for
Briefing reports5Onboarding a new volunteer to one stream
Data tables5Risk register, regulatory landscape, grants, strategic decisions, professional-advice queue
Slide decks5Community meeting, funding, legal, media kit, “what we know”
Audio discussions3Listening material — a synthesis, a “what we know”, and a critique of the project’s own weaknesses
Flashcard sets2Media training, and the open strategic choices

The audio is the strangest item and was genuinely useful: two synthetic voices discussing a community pub acquisition in Tasmania, generated from the group’s own documents, listenable on a drive. One of the three was deliberately commissioned as a critique of the project’s weaknesses — an easy thing to skip and the most valuable of the three.

What it was good for

Onboarding. A stream briefing that would have taken a volunteer an evening to write took minutes, and was accurate to the sources because it had nothing else to draw on.

Finding the shape of the corpus. The generated data tables surfaced cross-document patterns — a regulatory landscape assembled from six separate research memos — that nobody had assembled by hand.

Meeting material. Slide decks for a community meeting, generated rather than agonised over, freeing the time for the thing that actually mattered, which was the meeting.

What it was dangerous for

It is faithful to its sources, which means it inherits their errors. That is the entire risk in one sentence. When the corpus contained a debunked case study, or an over-confident figure, the generated briefing repeated it — cleanly, fluently, without the hedging that surrounded it in the original. Generated material launders uncertainty: a caveated paragraph in a working document becomes a crisp bullet in a briefing.

Its outputs are undated and unsourced by default. A briefing generated in May from a corpus that changed in June looks exactly like one generated in June. The workspace’s own header carried the warning, in the project’s standing formula: “All outputs are AI-generated and may contain errors — verify against primary sources before use in steering, legal, or public communications.”

Nothing generated from it was ever a source. It was onboarding material and meeting material. No claim entered the evidence register from a generated artefact; the protocol required a named, retrievable primary source, and a machine summary of the group’s own documents is not one.

Why none of it is published here

All twenty artefacts are withheld, and the reason is duplication rather than sensitivity. Twelve of them — the briefing reports, data tables and flashcard sets — are text files and are counted as withheld in the corpus map. The other eight are the slide decks and audio, which are binaries and sit outside that count; they are described in prose there instead.

Every one of them is generated from documents that are themselves in this archive. Publishing a machine’s briefing on the risk register beside the risk register adds a second, less reliable version of the same content, with the hedges stripped out — which is exactly the failure mode described above. The slide decks and audio carry the same problem in a form that is harder to check.

If you want what these artefacts contain, the sources are here, and they are better.

The transferable finding

Using a synthesis tool over your own corpus is genuinely useful for navigation and actively harmful as a source. The line the project drew — generated material can help a person find the document, and can never stand in for it — held for the whole campaign, and it is the line to draw.

End of document ← Back to the archive