No-Code/Low-Code Mechanistic Interpretability for AI Models

Modern AI systems, particularly complex architectures such as language models and multimodal models based on Transformers or diffusion architectures, exhibit a remarkable trend: they are becoming increasingly powerful and can autonomously handle ever more complex tasks, while at the same time their internal workings are becoming less and less transparent. This lack of transparency makes it difficult even for domain experts to trace and understand exactly how these systems arrive at their decisions and outputs, or to assess their emergent capabilities. This in turn hampers trust-building, safety validation, and ethical evaluation.

The emerging research field of Mechanistic Interpretability (MI) seeks to make these black-box systems transparent and comprehensible by investigating their internal structures using methods such as circuit analysis, activation engineering, parameter decomposition, and representation analysis. However, the methods employed require profound expertise in mathematics and computer science as well as deep knowledge of the field of generative AI – a significant barrier to entry for domain experts from psychology, ethics, the social sciences, and law. Yet their perspectives would be essential for interpreting MI results in order to develop robust and safe AI systems that are "human compatible" – that is, systems that place human preferences, norms, and values at their core and thereby enable socially responsible AI applications. This project investigates and develops innovative methods for integrating complex MI techniques into user-friendly No-Code/Low-Code (NC/LC) applications.

The core innovation lies in the novel synthesis of two disparate research fields: causal-analytical MI and democratizing NC/LC architecture. Particularly innovative is the exploration of multimodal and audio models – a research area that has so far been underrepresented in the MI context. The approach transforms the highly complex MI methodology through NC/LC principles and human-centered design in order to make it accessible for interdisciplinary research.

The project pursues five concrete scientific and technical work objectives:
(1) systematic analysis of the state of the art and selection of suitable MI methods, as well as extension of existing approaches,
(2) development of abstraction concepts and visualization paradigms,
(3) prototypical implementation of an interactive, web-based NC/LC MI platform,
(4) human-centered evaluation with representatives of the target groups, and
(5) application of the prototype to interdisciplinary research questions.

This fundamental research at the intersection of AI, cognitive science, and Human-Computer Interaction (HCI) is of high societal relevance: the tools developed catalyze an evidence-based public discourse, support the formulation of scientifically grounded AI policies, and foster collective trust in AI technologies through increased transparency. Democratizing AI analysis for diverse experts without AI expertise (e.g., in ethics and law) directly addresses the pressing societal challenge of algorithmic opacity and makes a substantial contribution to responsible AI development. 



Verbundprojektleitung


Projektleitung


Projektbearbeitung

Marc Guggenberger
T 09814877-541
marc.guggenberger[at]hs-ansbach.de

Project duration

2027-01-01 - 2030-12-31

Project partners

Hochschule Heilbronn
Northeastern University

Project funding

Bundesministerium für Bildung und Forschung

Funding programme

HAW-ForschungsAkzente