Concept-Targeted Attribution: New Framework Traces How AI Models Recognize Internal Concepts
Researchers introduced Concept-Targeted Attribution (CTA), a method that traces how AI models recognize concepts internally, even when those concepts aren't expressed in the model's output. The framework uses Cross-Layer Transcoders to build probe-specific circuits that explain why concept representations arise.