Archive

All 2,608 AI stories, newest first · page 90 of 109

Researchers Map 'Bias Fingerprints' in GPT-2 and Llama 3.2
research

Researchers Map 'Bias Fingerprints' in GPT-2 and Llama 3.2

A new study identifies specific neurons and attention heads in LLMs that encode harmful stereotypes. The research provides tools to locate and potentially mitigate biases in AI models. Researchers used a combination of contrastive neuron activation analysis and attention head tracking to pinpoint bias sources in GPT-2 Small and Llama 3.2.

via ArXiv cs.CL#llms#bias#stereotypes