Your Training Partner
Techniques Toolbox
Diagram of document analysis in five stages: inventory, triage on four criteria (relevant, current, authentic, credible), record in a six-column chart, reconcile, commit. The triage gate has a labelled reject exit, discard recording the reason, reconciliation fans out into three outcomes (duplicate, conflict, gap), and the gap loops back to the inventory as an arrow.

Document Analysis

Document analysis (BABOK 10.18) elicits business analysis information from the documents an organisation already holds: regulations, manuals, contracts, incident tickets, screens, data extracts, statutes and industry standards. The analyst selects the sources, appraises them, extracts findings and records them. It serves three uses BABOK distinguishes: understanding the context of a need, establishing how and why an existing solution works the way it does and checking against the written record what was said in an interview. Its artifact is a six-column chart, one row per finding, where the literal citation and the analyst's judgement occupy two separate columns and where the absence of a document is itself a finding. It is run before anyone else's time is spent, and its cost is controlled at triage.

Goal

Document analysis answers a question that comes before every other form of elicitation: what do we already know, in writing, before we start asking people. It draws from material the organisation already owns the information an analyst would otherwise have to chase through interviews, workshops and observation, and it does so without spending anyone's time. Its mirror question matters as much: what does the written record fail to say. Both answers govern the rest of the elicitation plan.

The technique produces four deliverables. The first is a documented picture of the current state, rules, procedures, data and constraints, which survives the departure of the business experts. The second is a traceable evidence base: every finding is attached to a source that is named, dated, versioned, citable and therefore defensible the day a stakeholder disputes it. The third, and the most profitable, is the list of questions worth asking, the ones the written record leaves open: a workshop prepared by a document analysis costs less and returns more. The fourth is a log of conflicts and gaps, which routes the remaining unknowns to the elicitation technique capable of resolving them.

Usage

When to use it

  • Solution to be replaced or extended: the technical and procedural documentation carries the current behaviour nobody remembers.
  • Business experts gone or unavailable: the written record is the only surviving witness of past decisions, and BABOK names this explicitly.
  • Regulated domain: statutes, ordinances and the collective labour agreement are sources of requirements.
  • Ahead of interviews and workshops: arriving informed shortens elicitation and raises the quality of the questions.
  • Interview finding to triangulate: a dated document confirms or contradicts what was said in the room.
  • Data requirements for a replacement system: forms, screens, reports and extracts give the fields, the formats and the volumes.
  • Steep hierarchy, low psychological safety: the technique forces nobody to contradict a senior stakeholder out loud.
  • Due diligence, audit, post-incident review: the written record is the object of the study itself.

When not to use it

  • Greenfield, future-state design: nothing existing to read, run workshops, brainstorming or prototyping.
  • Documentation known to be stale or never maintained: the source misleads, prefer observation and interviews.
  • Short, cheap decision: reading forty documents to settle two days of development is disproportionate, one interview is enough.

Description

The three uses of the technique

BABOK distinguishes three uses, and confusing them means missing the source that would have helped. Context research illuminates the environment of a business need: market studies, industry standards, org charts, internal memos, activity reports. Research on the existing solution establishes how a system works today and why it was built that way: business rules, technical documentation, training material, incident reports, prior requirements, procedure manuals. Validation checks an interview or observation finding against the written record, and this is triangulation: what a person reports in good faith and what the regulation prescribes diverge regularly, and the divergence is a finding in its own right.

BABOK further places data mining (10.14) among the approaches to document analysis. When the source is a dataset rather than a text, the technique keeps its nature and changes its instrument: this is its quantitative arm.

Running it in five moves

The technique runs as a filtering pipeline with exits. Five moves follow one another: inventory the candidate sources, triage those sources against four criteria, explicitly discarding the ones that fail, record the findings in a chart, reconcile the findings against one another to surface duplicates, conflicts and gaps, then commit the whole to the deliverables that will carry it. The two lateral exits of that pipeline are what makes the technique worth its cost: the reasoned rejection at triage, which contains the spend, and the referral to another elicitation technique when a conflict or a gap exceeds what the written record can settle.

Prepare: scoping and access to the sources

Preparation begins with one written sentence: the topics the analysis has to answer. Without that list the corpus has no edge, and the limitation BABOK names first, information overload, happens immediately. The topics are what entitles you not to read a document.

Then comes the inventory of candidate sources and their custodians, internal (regulations, manuals, incident tickets, prior requirements, screens, extracts, contracts) and external (statutes, ordinances, industry standards, collective labour agreements, market studies). Then obtaining access: rights, confidentiality and, as soon as personal data is involved, a legal basis for the processing. This is a cost line, routinely forgotten in estimates, and it is budgeted separately from the reading.

The ABPMP's BPM CBOK, which comes at the same ground from the process side, adds two source types to that inventory. Transaction and audit logs count among the written documentation to gather on a process, alongside its diagrams and whatever was produced when it was created. Audit reports belong there too, for the control points in the organisation they identify.

A project run under HERMES, the Swiss Confederation's project management method, gives that inventory a written starting point: every task description there follows a fixed structure, one rubric of which, the basics (Grundlagen), lists, where it is filled in, the outcomes already produced that the task needs as input. For the task they support, the analyst then holds a named list of the outcomes to gather, which they take up and complete with their own sources.

When the source is data rather than text, BABOK asks for two further definitions in preparation, and they decide what will be looked for. The data classes are the categories of information the question requires: identities, working time, absences, salaries, expense receipts. The data clusters are the elements grouped by a logical relationship: an employee file is a cluster, it gathers the identity, the contract, the time records and the expense claims. Naming the cluster before reading it is what shows that it is heterogeneous: its elements come from different sources and they do not obey the same retention rules. A cluster treated as one block produces a false requirement; the same cluster decomposed into classes produces the right one. This is also the junction with the quantitative arm: with no classes and no clusters defined, data mining (10.14) has no scope, exactly as reading has none without a list of topics.

Each candidate source is then appraised against four criteria, BABOK's, before a single line of it is read in detail. It is relevant if it bears on the scoped topics. It is current if it is still in force, which is judged on the date and the version. It is authentic if it is what it claims to be: an identified author, a provenance, an official version rather than a draft left in circulation. It is credible if it is accurate and sincere, the hardest criterion: a training manual describes the intended process, an incident report describes the actual process and both are sincere without saying the same thing.

Academic documentary research adds a fifth criterion BABOK leaves aside and practitioners pay dearly for: representativeness. The documents that reach you are the ones somebody kept. An archive of customer complaints is an archive of the written complaints, a procedure that was found is the one somebody thought worth keeping and the surviving corpus is no random sample of reality. A finding drawn from a document is true of that document; its generality remains to be established.

Triage produces an artifact, the source register, and it carries the reasons for rejection as much as the reasons for retention. Recording that "the 2016 training manual was not read, it describes a version of the system superseded in 2021" is a defensible finding; omitting it looks like an oversight.

Candidate sourceRelevantCurrentAuthenticCredibleDecision and reason
OLT 1 (SR 822.111), art. 73 and 73ayesyesyesyesRetained. Official consolidated text, a directly enforceable source of law.
CO (SR 220), art. 958fyesyesyesyesRetained. Same, for the accounting side of the employee file.
Collective labour agreement in forceyesyesyesyesRetained. A condition of applying art. 73a OLT 1: without it, the waiver of working-time recording is void.
Staff regulations, 2019 version (in force)yesyesyesqualifiedRetained with reservations. Paraphrases the law and departs from it. No requirement is drawn from it without checking against the legal text, and the departure is itself a finding.
Training manual for the current system, 2016yesnoyesqualifiedDiscarded. Describes a version of the system superseded in 2021. Reason recorded: a rejection is a finding.
2025 payroll extract (HRIS columns)yesyesyesyesRetained. Gives the actual state of the data, where the other sources give the documented state.
The source register: triage on BABOK's four criteria, before any detailed reading. The discarded row carries its reason, because an unrecorded rejection reads as an oversight.

Record: the document analysis chart

The review itself proceeds document by document, noting findings by topic, and it is recorded in a document analysis chart whose columns BABOK fixes: topic · type · source · verbatim excerpt · critique · follow-up. This is the technique's artifact.

One row per finding

A set of staff regulations produces eight findings across five topics: it occupies eight rows. A chart organised by document is read in the order of the sources, whereas the work is done in the order of the topics, and reconciliation becomes impossible.

The verbatim column carries the exact citation, with its article, page or version reference

It admits no rewording, and a cut inside a citation is marked with bracketed ellipses at the exact point of the cut. Six weeks later, when a manager disputes a finding, only a referenced citation survives the discussion: a paraphrase can be argued about indefinitely. This is the column that makes the analysis auditable. It is the first one damaged when time is short.

The critique column carries only the analyst's judgement

What the source implies, what it contradicts, what it omits, what it is worth. Merging the two columns destroys what the separation guarantees: that the reader can tell at a glance what the document says from what the analyst concludes. The follow-up column closes the row by making it actionable: requirement to write, conflict to settle with a named person, gap to carry into an interview.

Reconcile the findings against each other

The chart is sorted by topic, and three shapes appear. The duplicate: two sources say the same thing, they are merged keeping the authoritative source, and the other stays in the register. The conflict: two sources contradict each other, or a source contradicts what was said in an interview. A conflict is not settled by reading, it is triangulated, by questioning the author of the source or by observing the actual practice. The conflict between an internal document and the legal text it paraphrases is the most frequent and the most useful: it reveals a non-compliance nobody had seen.

BABOK names among the technique's limitations that the author of a document may not be available for questions. That is the exact counterpart of its strength: document analysis is often launched because the experts have left. Triangulating through the author presupposes an author, and the author is missing at the very moment the written record becomes indispensable. When neither the author nor the observable practice is reachable, the conflict is not settled: it is carried. The follow-up column earns its purpose here, transporting the open question, named, dated and attached to its two contradictory sources, to the body that will arbitrate it. A carried conflict is a known risk, which a steering group can decide to accept. A conflict settled at the desk by the analyst alone is a decision taken without a mandate, and it surfaces in production.

The gap is the third case, the technique's most poorly exploited finding. A scoped topic on which the corpus says nothing is information. It is written into the chart, with "no documentation found" as the source, and it is routed: further search, descent into sub-topics or referral to an interview, an observation or a workshop. A gap is found by holding the corpus against the list of topics: that is why the list of topics is written before the first reading.

Commit: from finding to deliverable

BABOK sets out two judgements at the moment the findings are poured into a work product. Do the content and level of detail suit the intended audience: the chart is a working instrument, and a steering group receives what it produces rather than the chart itself. And does the material gain from being turned into a visual: a graphic, a model, a process flow, a decision table. The audience also dictates one precaution: a finding that will contradict an executive's public account is validated with the author of the source before it is circulated, and it is then circulated as it stands.

The findings then migrate into the artifacts that will carry them for the long run: business rules (10.9), data dictionary (10.12) entries, data models (10.15) and data flow diagrams (10.13), decision tables, requirements. Each keeps a traceability link to the source document and its version. That link, and only that link, is what makes it possible a year later to answer "where does this requirement come from" and to revise it when the ordinance it rested on is amended.

What makes the exercise fail

Reading everything

With no topic scope, the corpus expands until the budget is exhausted, and the information overload BABOK names among the technique's limitations arrives before the findings do. The remedy sits upstream: the topics, written first.

Taking the document for the truth

Documentation gives the documented state, observation gives the practised state and the two diverge. The gap between the procedure manual and what the employee actually does is often the most valuable finding of the study, and it appears only on condition of going to look.

Paraphrasing in the verbatim column

The audit trail disappears and with it the ability to defend the finding. An abridged citation with no mark of the cut is worse still: it presents a rule that does not exist in that form, and the reader has no way of seeing it. This defect is painless on the day it is committed and expensive six weeks later.

Recording a finding without its reference or its version

It becomes unverifiable, therefore uninvestigable, therefore worthless. A requirement whose source is "the regulations" rather than "the staff regulations, 2019 version, art. 12 para. 3" cannot be re-checked.

Confusing "undocumented" with "non-existent"

An interface running in production that no document describes exists nonetheless, and the silence of the corpus says something about the organisation. The symmetrical error costs as much: a documented procedure is not thereby a followed procedure.

Anchoring on what exists

This is the limitation BABOK states, the technique mostly illuminates the current state, and the constraints of the system in place migrate silently into the requirements for the future one. The countermeasure is mechanical: mark every extracted constraint as inherited until somebody confirms it is still wanted.

The ABPMP's BPM CBOK adds an instruction for the reading: when working through a system's documentation, do not assume that the system in place is the best solution for the job. Two signals contradict it: users see the system as an obstacle to their work, or they have put in workarounds and manual steps to compensate for its shortcomings.

Skipping the regulatory sources

In a Swiss project, retention periods, the FADP and the collective labour agreement are requirements, and they are not negotiated in a workshop. An analysis that reads only the system's documentation misses the requirements that cannot be waived.

AI considerations

Of all the BABOK techniques, document analysis is the one language models change most: its dominant cost is reading a large corpus, which the machine does very well, and its dominant failure is believing the document, which the machine does very badly.

Triaging a large corpus is the most profitable use. Classifying and clustering the sources by topic before any human opens one attacks head-on the limitation BABOK names, overload, and moves the human effort towards the documents that carry the scoped topics. Pre-filling the chart follows: a model can populate topic, type, source and, under an explicit instruction to quote without rewording, verbatim. The critique column stays with the analyst, because it is the judgement, and judgement is not delegated. Detecting duplicates and conflicts by semantic similarity surfaces, in minutes and across hundreds of documents, the near-identical passages and the contradictory statements: retrieval is a model's strength, arbitration is not. A multilingual corpus is an ordinary Swiss condition, staff regulations in German, a collective labour agreement in French, a vendor's documentation in English: machine translation serves the reading, and the verbatim cell stays in the language of the source. Optical character recognition, finally, makes a scanned archive searchable, which changes the order of magnitude of the corpus within reach.

Four limits hold, and there is no way round them. Currency and authenticity are metadata work: no model knows that the 2016 training manual describes a version of the system superseded in 2021, because that information is not in the document, it is beside it. This is why BABOK places those criteria in preparation. Credibility escapes the model, which reads every document as if it spoke true: telling the intended process of a training manual from the actual process of an incident report requires knowing what a training manual is. The hallucinated verbatim is the gravest risk, because a model produces a fluent, plausible citation that is not in the document or one that elides a condition without flagging it: every verbatim cell is re-checked against the source, reference in hand, and a chart of unverified citations is worth less than no chart at all, since the reader trusts it. Then silence: a model summarises what it was given and cannot report what it was not given. Gaps are found by holding the corpus against the list of topics, which is an act of analysis.

That leaves confidentiality, which is a first-order constraint here. The documents worth analysing are exactly the ones that contain personal data: payroll extracts, HR files, contracts, patient records. The FADP (SR 235.1, in force since 1 September 2023) applies, and the FDPIC has said so: because the act is technology-neutral, it applies directly to processing carried out by means of artificial intelligence, without any need for a specific statute. A public access point to a model is a disclosure to a third party. The work is run on an instance whose hosting location is known and which is contractually barred from training its models on the inputs. Failing that, the corpus is redacted before ingestion.

Examples

A manufacturing SME in French-speaking Switzerland, 240 employees, is replacing its time and absence management system. Before the first workshop, the business analyst runs document analysis on the staff regulations, the industry collective labour agreement, the screens and spreadsheets in use, the payroll extract and the applicable federal law. Here is the filled-in chart, in BABOK's six columns.

TopicTypeSourceVerbatim excerptCritiqueFollow-up
Retention period for time recordsFederal ordinanceOLT 1 (SR 822.111), art. 73 para. 2"Les registres et autres pièces sont conservés pendant un minimum de cinq ans à partir de l'expiration de leur validité."Firm legal constraint: records are kept for at least five years from the expiry of their validity. The staff regulations announce three years, so the current system purges too early.Requirement: retain for at least five years. Discrepancy with the regulations to be settled with HR.
Retention of accounting records (expense claims)Federal actCO (SR 220), art. 958f para. 1"Les livres et les pièces comptables ainsi que le rapport de gestion et le rapport de révision sont conservés pendant dix ans. Ce délai court à partir de la fin de l'exercice."Ten years from the end of the financial year, so a different period from the time records. The employee file is a heterogeneous cluster: two purge rules coexist inside it.Requirement: purge by category of record, never by file.
Waiver of working-time recordingFederal ordinanceOLT 1, art. 73a paras. 1 and 2Para. 1: "Les partenaires sociaux peuvent, dans une convention collective de travail (CCT), prévoir que les registres et pièces ne contiennent pas les données prévues par l'art. 73, al. 1, let. c à e et h, si les travailleurs concernés:
a. disposent d'une grande autonomie dans leur travail et peuvent dans la majorité des cas fixer eux-mêmes leurs horaires de travail;
b. touchent un salaire annuel brut dépassant 120 000 francs (bonus compris) ou la part correspondante en cas de travail à temps partiel, et
c. ont convenu individuellement par écrit de renoncer à l'enregistrement de la durée du travail."
Para. 2: "Le montant du salaire annuel brut visé à l'al. 1, let. b, est adapté à l'évolution du montant maximum du gain assuré LAA."
Four cumulative conditions: a collective labour agreement that provides for it (the opening clause), a high degree of autonomy with the ability to set one's own hours in the majority of cases (let. a), a gross annual salary above CHF 120'000 including bonus (let. b) and an individual written agreement (let. c). The 2019 staff regulations cite only two of them, the salary and the written agreement: they omit the collective agreement and the autonomy. Without a collective agreement the waiver is void, and the company is currently applying an exception that does not exist. The threshold is indexed by para. 2 to the maximum insured earnings under the Accident Insurance Act (LAA): it is not a constant.Conflict between the internal regulations and the ordinance. Business rule to write (10.9) carrying all four conditions, to be validated with the head of HR. The threshold is written as a reference to the LAA indexation, never as a hard-coded number in the system.
Format of the AVS numberData extract2025 payroll extract, column NAVS13756.1234.5678.97Thirteen digits, format 756.NNNN.NNNN.NC, last digit a check digit. No format validation in the current system.Format constraint to carry into the data dictionary (10.12).
Salary declaration to the compensation fundGapNo documentation foundNo excerpt: the source is absent.Gap. The Swissdec interface (the unified salary declaration procedure) is in production, and no document describes it.Gap. Elicit by interview with the vendor. Do not infer the interface from the code.
A filled-in document analysis chart, five findings, four sources and one absence of source. The verbatim excerpt and critique columns stay apart: the first is checked against the source, the second commits the analyst. The verbatim cells stay in the language of the source, French, because OLT 1 and the CO have no English version and a translated citation stops being a citation.

Five rows, five different lessons, and this is why a chart is read by topic. The first carries a firm legal constraint the current system violates. The second reveals a conflict between two legal retention periods living side by side in the same employee file, and the heterogeneous cluster is what is paid for here: an apparently simple purge requirement becomes a requirement to purge by category of record. The third is a conflict between an internal document and the law it paraphrases: the staff regulations omit two of the ordinance's four cumulative conditions, including the collective agreement without which the exception does not exist, and document analysis surfaces the non-compliance with no interview having taken place. It also teaches how to write an indexed threshold: as a reference to the text that indexes it. The fourth is a format finding drawn from a data extract, and it goes straight into the data dictionary. The fifth cites nothing, because its source is the absence of a source: it is a gap, it is a finding like the others and it exits the technique towards an interview.

Visualisations

The technique leaves two artifacts, and each imposes its own form. The run is a spatial object: a chain of stages, a triage gate with a lateral exit, a return loop from the gap back to the inventory, a fan of destinations at the end. What has to be seen are the exits, and a numbered five-point list would destroy them, suggesting a linear sequence you enter at one end and leave at the other. It is drawn. The chart falls under a different rule: six columns, one row per finding, text in every cell, therefore a table, selectable, searchable and legible on a phone. The source register follows the chart, for the same reasons.

The pipeline also serves as a control on an analysis in progress. Three questions apply to it. Does the triage gate have a rejection exit that is used, with written reasons, or were all the candidate sources read, in which case there was no triage and the budget is already committed. Do the gaps loop back to the inventory, or did the analysis merely summarise what it happened to find. Do the findings exit towards deliverables traced to their source, or is the chart its own terminus, in which case it will die with the project.

Cost

PhaseLevelJustification
PreparationMediumScoping the topics, defining the data classes and clusters, inventorying the sources and above all obtaining access: rights, confidentiality, a legal basis for processing when personal data is involved. Cheap where the repository is indexed and searchable, expensive where the archive is on paper, in silos or in legacy formats.
ExecutionHighReading grows linearly with the number of sources, and BABOK names the failure: a wide range of sources makes the effort very time-consuming and produces overload. This is the dominant line and the one estimates underrate.
DocumentationMediumThe chart is filled in while reading, so most of the documentation runs concurrently with execution. The residual cost is real but bounded: turning findings into deliverables and holding the traceability to the source and its version.

The asymmetry is the striking fact in that table: document analysis is cheap to start and expensive to finish. Nothing is easier than opening a first document, and the total cost is decided at the scoping of the topics and the triage of the sources, that is, in the phase that looks the least productive.

Tooling

A spreadsheet, or a table in the team wiki, is enough to carry the chart and the register, as most analyses do. The artifact has six columns: no specialised tool is required to write it, and a specialised tool that discouraged anyone from keeping it would be a regression.

Document management systems (SharePoint, Confluence, an enterprise electronic document management system) carry the corpus. Where the organisation runs a genuine records management system compliant with ISO 15489-1:2016, the source register is half built in advance: the standard requires the metadata of context, content, structure and management over time that make a source appraisable and therefore datable, authenticable and versionable.

Full-text search and optical character recognition become indispensable past a certain corpus size. A search engine over the repository, or ripgrep over a local export, and an OCR engine (Tesseract, ABBYY FineReader) for scanned documents. An unindexed corpus is paid for twice, in search time before it is paid for in reading time, and that is the hidden cost of the Execution line.

Reference managers (Zotero) hold source metadata and the citation, which is the source register under another name, and they suit a largely external corpus well: statutes, ordinances, standards, studies.

Qualitative coding software (NVivo, MAXQDA, ATLAS.ti or Taguette as open source) is the category that turns document analysis from a sustained read into a method. Passages are coded by topic across the whole corpus, and the chart is filled by querying the codes rather than by retyping. The threshold is empirical: as soon as re-reading by topic costs more than the initial reading, coding pays for itself.

The quantitative arm uses different instruments. When the sources are data rather than texts, SQL, a data profiler or OpenRefine replace reading with sampling, and the technique hands over to data mining (10.14) and the data dictionary (10.12).

Requirements tools (Jira, Polarion, ReqView, DOORS) serve the technique's output: the link from the requirement to the source document and its version. That is the deliverable which survives the project.

The LLM and retrieval-augmented stack, finally, under the constraints of the FADP: an instance whose processing location is known and which does not train its models on the inputs, as soon as personal data is in the corpus.

Sources

Desk Check
All techniques
Dot voting