Sign InOpen Brain
arXivPaperNeeds Review

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

A controlled study encoded authenticated internal-state evidence into unchanged answers, suggesting generated text could carry provenance signals, but not that current models reveal them naturally.

arXiv · Aug 17, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Researchers forced arithmetic models through one of two discrete internal paths, authenticated the chosen state, and encoded a subtle detectable pattern in otherwise equivalent output. Feed-forward and transformer systems passed **all 128 matched pairs** in public and sealed evaluations.

Practical Implication

For builders, this sketches a future provenance mechanism in which output carries evidence about a causally relevant computation, not just the final answer. Replication across **five feed-forward models** and **three transformers** supports the controlled mechanism.

Agent-Ready Context
Researchers forced arithmetic models through one of two discrete internal paths, authenticated the chosen state, and encoded a subtle detectable pattern in otherwise equivalent output. Feed-forward and transformer systems passed **all 128 matched pairs** in public and sealed evaluations.

For builders, this sketches a future provenance mechanism in which output carries evidence about a causally relevant computation, not just the final answer. Replication across **five feed-forward models** and **three transformers** supports the controlled mechanism.

This is a bounded proof of concept using deliberately trained architectures and an engineered signal. In a separate answer-only transformer experiment, linear probes did not recover a naturally learned intermediate state, so the work does not validate provenance for ordinary model outputs.
Connected Context · Feed7 Judgment

This narrows provenance claims from inspecting outputs or probing naturally learned representations to deliberately training an authenticated causal state into the output. The controlled success shows that computation-path evidence can be carried, while the failed answer-only probe warns that ordinary models cannot yet be assumed to expose such provenance naturally.

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard ModelsThe compliance audit shows probes can appear accurate without reading the governing rule; this work contrasts that failure with an engineered signal tied to a forced, authenticated internal path.QuoteBench: How Matched Scores Can Hide Command-Path FailuresQuoteBench shows final scores can conceal which command path produced them; computational provenance sketches a complementary way to attach evidence about a causally relevant internal path to the output itself.The Low Frequency Trap: Video Language Models Fail at Simple Event BookkeepingBoth reject answer-only confidence: the video study requires event-level traces, while this work tests whether evidence of an intermediate computational state can travel with the answer.
Context Map
benchmarksecurity#agent-evals#benchmark-integrity
Uncertainty
This is a bounded proof of concept using deliberately trained architectures and an engineered signal. In a separate answer-only transformer experiment, linear probes did not recover a naturally learned intermediate state, so the work does not validate provenance for ordinary model outputs.