Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text
Hiding proof of how an AI reached its answer inside the answer itself
Researchers embedded hidden fingerprints into text generated by AI models that reveal which internal computational path the model used to reach its answer—even when different paths produce identical outputs. In controlled tests with both simple neural networks and transformers, a detector could later read these fingerprints and identify the verified internal state that was actually used, succeeding on all 128 test cases.
Today's AI systems are black boxes: you see the answer but not how the model produced it. If AI systems could cryptographically prove which computational steps they took, it would enable auditing, safety verification, and accountability—critical for high-stakes applications like medical diagnosis or financial decisions where knowing the reasoning matters as much as the result.