Skip to main content
All resources

Articles

Evidence-grounded AI: what proposal work should demand of a model

What source traceability actually requires when an AI-drafted claim carries contractual weight — and why citing a source is not the same as governing it.

Why "a human reviews it" is not a control

Almost every vendor in this market answers the hallucination question the same way: the model drafts, a person checks. As a description of the workflow that is fine. As a governance position it is weak, because it puts the entire control on the least reliable point in the process — a reviewer reading fluent, plausible text under deadline, looking for the one sentence that is confidently wrong. Fluency is precisely what makes that hard. The reviewer is not verifying claims; they are reading prose and forming an impression, and an unsupported sentence that reads well passes.

A citation proves the sentence came from somewhere. It does not prove that somewhere was right, current, or approved by anyone. In proposal work those are the parts that carry liability.

The vocabulary already exists, in a literature nobody here is citing

This problem has been worked on seriously, just not by proposal vendors. Governance frameworks for AI systems treat citations and source traceability as a named control rather than a feature — the FINOS AI Governance Framework, for example, sets out providing citations and source traceability for AI-generated information as a mitigation whose purpose is to make claims auditable and to expose reliance on outdated or inappropriate sources. Alongside it there is a growing research literature on claim-level grounding and inline citation generation: attaching each individual assertion to the specific passage that supports it, rather than listing documents that were consulted. That distinction — per-claim grounding versus per-document attribution — is the one that matters in a proposal, and it is largely absent from how the tooling in this market is described.

The proposal-specific version of the control

Applied to a submission, grounding has to answer four questions about every assertion, not one. Most tools answer the first and stop.

  1. Where did this come from? A specific record or page, not a document list. "Sourced from the project archive" is not traceability.
  2. Is that source current? A completion certificate does not expire, but a certification does, a methodology is superseded, and a reference contact changes employer. A citation to a stale record is worse than no citation, because it looks like diligence.
  3. Is that source approved? Somebody with the authority to do so has to have said this is the version the firm stands behind. An unapproved file that happens to be in the repository is not a source; it is a draft somebody left there.
  4. What is the claim committing us to? A statement about past experience is a factual claim. A statement about an expert's availability is a contractual undertaking. They need different levels of scrutiny, and only a person can tell you which one a sentence is.
Our reading, not a rule
The four-question test is ProposalOS's framing, built on the citation and traceability controls the AI-governance literature already describes. It is not a published standard.

Traceability without governance is theatre

This is the part the market is currently getting wrong while sounding right. Source attribution has become the headline differentiator in proposal tooling, and a system that cites its sources is genuinely better than one that does not. But if the underlying library is ungoverned — no owners, no approval state, no review dates — then citation does not fix anything. It makes an unreliable claim look verified, and it does so at the exact moment a reviewer is deciding whether to trust the output. The governance of the source is not a separate concern from the citation; it is what makes the citation mean anything.

The tasks where assistance genuinely pays for itself

Reading long tender documents, restating a clause as a tracked requirement, finding the past projects most like this one, checking a draft against a requirement list, spotting what has not been answered — these are high-volume, well-specified and, crucially, verifiable. A person can confirm the output faster than they could have produced it, which is the whole test. Note that all of these are retrieval and checking rather than composition. That is not an accident: retrieval is where the hours actually go, and it is also where a wrong answer is cheapest to catch.

The decisions that must stay with people

Whether to bid. Whether a reference is genuinely comparable. Whether a commitment is one the firm can actually meet. Whether the submission is ready to go. These are judgement and liability, and neither transfers to a tool. A system that quietly makes them has not saved effort — it has moved risk to somewhere nobody is watching.

Three claims that should never be model-generated

Beyond the general division of labour, some assertions carry consequences that make generation inappropriate regardless of how good the grounding is.

  • Eligibility statements. Whether you meet a registration, turnover or classification threshold is a fact about your organization with disqualification attached to it, and it should be asserted by someone who can be held to it.
  • Certifications and their validity. A model can retrieve the certificate; it should not characterise its status. Expiry is a date, and the answer is either current or it is not.
  • What a named expert is committed to. A CV in a formal tender is an undertaking about a specific person's availability and qualifications. That person, and whoever signs the submission, are the only appropriate authors of it.

How to evaluate a vendor's traceability claim

Nearly every platform in this market now says it cites sources. The claims are not equivalent, and the differences are easy to test in a demonstration if you know what to ask.

  • Ask it to show the source for one specific sentence, not for the section. If the answer is a list of documents, that is document-level attribution, not claim-level grounding.
  • Ask what happens when it cannot support a statement. The useful answer is that it marks the gap; the answer to be wary of is that it writes something reasonable.
  • Ask whether the cited record has an owner, a version and an approval state, and whether the interface shows them at the point of citation.
  • Ask what it does with a source that is out of date. Silence here is the common case, and it is the failure mode that reaches an evaluator.
  • Ask which actions it will not take without a person — insertion, approval, export, submission — and confirm that in the product rather than in the brochure.

Primary sources

Quotations are reproduced from the published document at the date shown above. Procurement documents are revised, and the request for proposals you receive governs your bid — check the clause in your own tender before relying on it.