Go directly to page contentent
TEAM LADE
Diego Doimo
Researcher at Laboratory of Data Engineering

 

I am a researcher at the Laboratory of Data Engineering at Area Science Park and a contract professor of Natural Language Processing at the University of Trieste.

 

My research focuses on understanding how language and vision-language models
represent and process information. I am interested in white-box interpretability methods to study how knowledge is organized and transformed inside deep neural networks, analyzing the relationship between the geometry and semantics of neural representations. I also study how to combine mechanistic interpretability with geometric and unsupervised methods from manifold learning to characterize the computational principles underlying modern multimodal models.

 

I got my Ph.D. from the International School for Advanced Studies (SISSA, Trieste), where I developed an algorithm for estimating the intrinsic dimension of high-dimensional datasets. Using this method, I showed how geometric compression is linked to semantic abstraction in transformers trained with self-supervision. I also investigated the density structures of hidden representations in convolutional neural networks, showing how they relate to the hierarchical organization of concepts, and how the low-dimensional geometry of hidden representations can explain the generalization capabilities of overparameterized neural networks.

 

Research Interests

 

  • Geometric methods for the interpretability of neural representations
  • Mechanistic interpretability of language and vision-language models
Experience & Education
  • Ph.D. in Physics and Chemistry of Biological Systems (SISSA)
  • Master in Physics of Complex Systems (PoliTO & Sorbonne Université)
Latest Publications
02/06/2026
Visual Instruction Tuning Aligns Modalities through Abstraction
Abstract Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image…
Go to the news Visual Instruction Tuning Aligns Modalities through Abstraction
18/09/2025
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
Abstract Recent advances in multimodal training have significantly improved the integration of image understanding and…
Go to the news The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
10/12/2024
The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language Models
Abstract In-context learning (ICL) and supervised fine-tuning (SFT) are two common strategies for improving the…
Go to the news The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language Models