Mapping Moral Reasoning Circuits: A Mechanistic Analysis of Ethical Decision-Making in Large Language Models

Abstract

This paper systematically investigates how large language models (LLMs) encode moral reasoning across six moral dimensions: care, fairness, loyalty, authority, sanctity, and liberty. We propose a novel interpretability pipeline that combines differential activation analysis, automated neuron description, and ablation experiments to identify specialized neurons aligned with each moral dimension. Our curated dataset of 240 validated moral and immoral statement pairs guides this exploration and reveals that certain neurons consistently exhibit increased activation in response to morally aligned statements. Notably, the care and sanctity dimensions show the largest sets of specialized neurons, whereas fairness and loyalty show fewer. We further demonstrate that ablating these neurons can causally modulate ethical decision-making, supporting the presence of discrete sub-circuits that influence moral outputs. Our findings not only advance the theoretical understanding of moral reasoning in LLMs, but also highlight avenues for targeted interventions and alignment. 

mehr

Mehr zum Titel

Titel Mapping Moral Reasoning Circuits: A Mechanistic Analysis of Ethical Decision-Making in Large Language Models
Medien In: Degen, H., Ntoa, S. (eds) Artificial Intelligence in HCI. HCII 2025. Lecture Notes in Computer Science, Springer, Cham
Verlag Springer Nature Switzerland
Herausgeber Degen, Helmut; Ntoa, Stavroula
Band 15820
ISBN 978-3-031-93415-5
Verfasser Prof. Dr. Sigurd Schacht, Carsten Lanquillon
Seiten 97-116
Veröffentlichungsdatum 01.06.2025
Projekttitel TTZ NEA (hoheitlich)
Zitation Schacht, Sigurd; Lanquillon, Carsten (2025): Mapping Moral Reasoning Circuits: A Mechanistic Analysis of Ethical Decision-Making in Large Language Models. In: Degen, H., Ntoa, S. (eds) Artificial Intelligence in HCI. HCII 2025. Lecture Notes in Computer Science, Springer, Cham 15820, 97-116. DOI: 10.1007/978-3-031-93415-5_6