Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Circuit Localization in a Tiny Language Model: Geographic Routing and Representational Depth in Qwen2-0.5B

Domaine:

natural language processing

Type de record:

model
Créateur:
Ela
Éditeur:
Zenodo
Hôte:avatar

Abstract This paper investigates how the Qwen2-0.5B language model encodes and routes geographic associations. Using mechanistic interpretability techniques, including activation patching, causal ablation, and head-level path patching, we localize the circuits responsible for geographic fact recall.

Key Findings:

  • Mechanistic Bottleneck: We identify a critical "Contextual Integration Hub" at Layer 17. Specifically, patching out Layer 17 Head 7 alone causes a 103% damage to the model's geographic prediction, flipping the logit preference to a corrupted target.

  • Representational Depth Disparity: We find a significant difference in how the model processes Western vs. MENA (Middle East and North Africa) geography. Western concepts resolve significantly earlier (mean Layer 18), while MENA concepts are consistently gated to the deepest layers of the model (Layer 21), regardless of prompt specificity.

  • Routing Convergence: High-specificity cues (landmarks/currencies) trigger earlier circuit activation than ambiguous cultural cues, suggesting a multi-mode routing architecture based on cue strength.

Methodology: Experiments were conducted using the TransformerLens library. The study includes:

  1. Multi-trigger convergence tests (Economic vs. Landmark vs. Cultural cues).

  2. A large-scale geographic bias audit (36 concepts).

  3. Layerwise causal ablation sweeps.

  4. Precision head-level path patching for circuit localization.

Code and Reproducibility: The full codebase, including Jupyter notebooks for all experiments (core experiments and diagnostics), is available on GitHub: github.com

Visit

doi.org

Tasks

language modeling

Tags

Mechanistic InterpretabilityQwen2Circuit LocalizationActivation PatchingCausal InterventionRepresentational DepthAI SafetyTransformerLensLarge Language ModelsGeographic Bias

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Tiny Aya: Bridging Scale and Multilingual DepthianMuchesia/tiny-translation-modelKeynes’ Finance Circuit model on banks in AfricaHAUSA LANGUAGE TYPOLOGY: AN IN-DEPTH STUDYBilingual Representational Systems in Free RecallPP-OCRv6 (tiny) multilingual handwritten and printed text recognition model

Tiny Aya: Bridging Scale and Multilingual Depth

Tiny Aya redefines what a small multilingual language model can achieve. Trained on 70 languages and

ianMuchesia/tiny-translation-model

Minimal seq2seq Transformer for English->Swahili translation with attention-based encoder/decoder bl

Keynes’ Finance Circuit model on banks in Africa

Abstract Since the publication of Keynes General Theory in 1936 when Keynes developed an original F

HAUSA LANGUAGE TYPOLOGY: AN IN-DEPTH STUDY

This paper presents a typological account of Hausa, an Afroasiatic language spo

Bilingual Representational Systems in Free Recall

40 undergraduates equally proficient in English and Yoruba were classified as “low” or “high” on sep

PP-OCRv6 (tiny) multilingual handwritten and printed text recognition model

PP-OCRv6 (tiny) multilingual text recognition base model Description This is the tiny variant (~0.