← Research archive
Experiment verified Sep 4, 2026

Claim-Level Verification and Clean Corpus Consolidation

KESPA consolidated strict claim-level verifier survivors from a 100-card diagnostic subset into a clean 423-claim corpus. The experiment recorded a 48.0% strict claim survival rate and rescued 331 independently supportable claims from cards that would have been discarded under whole-card acceptance.

claim-verification knowledge-quality corpus-consolidation atomic-claims provenance
MARKDOWN

README.md

2,300 bytes SHA-256 5e0b7548eb8f9d86…

KESPA-EXP-005 — Claim-Level Verification and Clean Corpus Consolidation

Date: 2026-09-04 Status: Verified — user environment Project: KESPA AI / NexLabs Studios

Research question

Does strict claim-level verification preserve independently supportable knowledge that would be lost if generated knowledge is accepted or rejected only at whole-card granularity?

Background

KESPA's earlier generation pipeline produced multi-claim cards.

That creates a trust-boundary problem: one unsupported claim can make the whole card unsafe to trust, even when several other claims in that same card are independently supported by frozen evidence.

The diagnostic experiment therefore moved the verification boundary from:

card -> accept/reject

to:

claim -> evidence -> strict verification -> retain/reject

Result

The clean consolidation produced:

  • 423 verified atomic claims
  • from 66 source cards
  • spanning 66 topics
  • linked to 56 frozen evidence excerpts
  • 331 claims rescued from cards that otherwise failed
  • recorded strict claim survival: 48.0%

The 331 rescued claims represent about 78.3% of the final 423-claim corpus.

Why this matters

A whole-card verifier can throw away valid knowledge simply because a neighboring claim fails.

Claim-level verification gave KESPA a finer trust boundary:

  • unsupported claims could be rejected;
  • independently supportable claims could survive;
  • each surviving claim remained linked to frozen evidence/provenance.

That 423-claim corpus later became the clean baseline used for the retrieval and A/B/C benchmark work.

Safety / production boundary

This consolidation step did not:

  • modify production Chroma;
  • modify Brain API behavior;
  • write to the production database;
  • assign production trust;
  • mark the corpus production-ready.

The historical result was a clean candidate corpus suitable for controlled benchmarking and later promotion decisions.

Limitation

The diagnostic set covered 100 generated cards, not the full 1,075-candidate generation population.

The result therefore supports claim-level verification as the better trust boundary for this diagnostic subset; it does not establish a universal survival rate for all future KESPA knowledge.

TEXT

SHA256SUMS.txt

399 bytes SHA-256 7320b6baa3def64f…
5e0b7548eb8f9d861fa2adbb666bb3da0e367f3a507fed75f91298a7dc67f10a  README.md
42606fa5c57223de9b7e77d68ab680c2ff20457c63d717549e55b087e1eb07f0  experiment.json
0bcb3b72d67e14b8bd2f5413d007be8b602a8f377ad36e71d6dbd8260421b25c  results.csv
1bbc1c7a09d1daa1bea98e2ab053375740cb85add5b234e2cde8252079845b12  methodology.md
973305e931970ac71c50d83b5493eeeae76c3af2e5c35340be119998aaab06f4  provenance.json
JSON

experiment.json

2,096 bytes SHA-256 42606fa5c57223de…
{
    "schema": "kespa.public_experiment.v1",
    "id": "KESPA-EXP-005",
    "title": "Claim-Level Verification and Clean Corpus Consolidation",
    "date": "2026-09-04",
    "status": "verified",
    "research_question": "Does strict claim-level verification preserve independently supportable knowledge that would be lost if KESPA accepted or rejected generated knowledge only at whole-card granularity?",
    "diagnostic_scope": {
        "source_cards": 100,
        "verified_corpus_source_cards": 66,
        "topics_represented": 66
    },
    "result": {
        "verified_atomic_claims": 423,
        "linked_evidence_excerpts": 56,
        "strict_claim_survival_percent": 48,
        "claims_rescued_from_failed_cards": 331,
        "production_ready": false
    },
    "derived_measurements": {
        "evaluated_claims_approximate_from_reported_survival_rate": 881,
        "rescued_claim_share_of_final_corpus_percent": 78.3
    },
    "interpretation": [
        "Whole-card acceptance was too coarse a trust boundary for the diagnostic candidate set.",
        "Claim-level verification allowed KESPA to retain individually supportable claims even when another claim in the same generated card caused the card as a whole to fail.",
        "The resulting 423-claim corpus became the clean verified knowledge baseline used by subsequent retrieval and A/B/C experiments."
    ],
    "safety_and_data_impact": {
        "production_chroma_changed": false,
        "brain_api_changed": false,
        "database_changed": false,
        "trust_assignment_changed": false,
        "production_ready_changed": false
    },
    "limitations": [
        "The clean corpus covered the 100-card diagnostic subset, not all 1,075 generated candidates.",
        "The reported 48.0% value is the recorded strict claim survival rate from the historical experiment; the public package does not reconstruct omitted raw candidate rows.",
        "Surviving verification means the claim met the experiment's strict evidence-verification rules; it is not a universal or permanent truth guarantee.",
        "No production trust assignment or production-ready status was granted by this consolidation step alone."
    ]
}
MARKDOWN

methodology.md

986 bytes SHA-256 1bbc1c7a09d1daa1…

Methodology

Input population

The historical experiment used the existing scale claim/evidence verifier against a 100-card diagnostic subset of generated knowledge.

Evidence resolution used canonical URL plus SHA-256 rather than URL alone so that multiple legitimate frozen excerpts from the same authoritative URL could remain distinct.

Verification unit

Verification was performed at the atomic-claim level rather than the whole-card level.

Each surviving claim retained evidence/provenance linkage. Claims that did not meet the strict support criteria were excluded from the clean corpus.

Consolidation

The verified survivors were consolidated deterministically into:

clean_verified_claim_corpus_v001

The historical user-environment validation recorded the exact expected count of 423 claims.

Production isolation

The consolidation did not write to production Chroma, Brain, or the database, and did not assign production trust or production-ready status.

JSON

provenance.json

1,517 bytes SHA-256 973305e931970ac7…
{
    "schema": "kespa.public_provenance.v1",
    "research_id": "KESPA-EXP-005",
    "source_basis": "Historical KESPA changelog/Jira export supplied for public research reconstruction.",
    "recorded_artifact_hashes": [
        {
            "path": "verified_corpus/clean_verified_claim_corpus_v001/manifest.json",
            "recorded_sha256": "48764DDEDB2D094DCD621E030321653C557494C872EE9ADAC405FD58022FAAA9"
        },
        {
            "path": "verified_corpus/clean_verified_claim_corpus_v001/validation_summary.txt",
            "recorded_sha256": "30F75EC56D22C811CFC171F31F1689AF7126C3DE6E9B12F9108EA672C77A91EA"
        },
        {
            "path": "verified_corpus/clean_verified_claim_corpus_v001/verified_claims.jsonl",
            "recorded_sha256": "D0BD8E02484472263A8D38E82023DD8CBEFC4B70AC86BA31713E04A63044352B"
        },
        {
            "path": "verified_corpus/clean_verified_claim_corpus_v001/verified_evidence.jsonl",
            "recorded_sha256": "0C181ED53224EE5E14A09037A49840F3D3A171C7CFF8322E090C3C64ABF13C4A"
        }
    ],
    "hash_verification_note": "These SHA-256 values are transcribed from the historical user-environment record. The underlying clean-corpus files were not supplied in the current attachment, so this public package does not claim to have independently rehashed them.",
    "publication_note": "The public package publishes aggregate corpus-consolidation measurements, methodology, limitations, and recorded provenance hashes. The full verified-claims/evidence JSONL data and private implementation scripts are not republished here."
}
CSV

results.csv

411 bytes SHA-256 0bcb3b72d67e14b8…
metric,value,unit_or_status
diagnostic_source_cards,100,cards
final_verified_source_cards,66,cards
topics_represented,66,topics
verified_atomic_claims,423,claims
linked_evidence_excerpts,56,evidence_excerpts
strict_claim_survival,48.0,percent
claims_rescued_from_failed_cards,331,claims
rescued_claim_share_of_final_corpus,78.3,percent
full_generation_population,1075,candidate_cards
production_ready,NO,status