Go directly to page contentent

Scientific Publications

04/04/2026

Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective

Abstract A key challenge in machine learning is to explain how learning dynamics select among the many solutions that achieve identical loss values in overparameterized models—a phenomenon known as implicit bias. Controlling this bias provides a direct mechanism on learned representations, which are central to interpretability, robustness, and reasoning in modern AI systems. Yet, despite its importance, existing explanations remain largely ad hoc and lack a unifying mechanism. We develop a theoretical and constructive framework in which implicit bias emerges as a geometric correction induced by the interplay between gradient noise and continuous symmetries of the loss. We compute the induced bias across a range of architectures, predicting new behaviors and explaining known ones. The approach also enables inverse design: by engineering predictorpreserving parameterizations, it is possible to shape the bias, with sparsity and spectral sparsity emerging as canonical instances. Numerical experiments support the theory and validate the inverse-design framework in controlled settings. Authors Nicola Aladrah, Emanuele Ballarin, Matteo Biagetti, Alessio Ansuini, Alberto d’Onofrio, Fabio Anselmi Journal Preprint Publication Date 04/04/2026 Consult the publication

11/03/2026

Chirality Unlocks Sub-Terahertz Molecular Dynamics in Peptide Nanotubes

Abstract Low-frequency modes in the picosecond range play a key role in molecular recognition and biomolecular function. Here, we investigate how these dynamics at the interfaces of self-assembled peptide biomaterials relate to fine structural details of the supramolecular assembly. We resolve terahertz and mid-infrared modes in individual diphenylalanine nanotubes by nanoscale microspectroscopy, uncovering a direct correlation between supramolecular chirality, molecular arrangement, and picosecond dynamics at the nanotube-environment boundary. By combining THz and mid-IR nanoscopy with a multiscale strategy spanning single nanotubes and ensemble measurements, supported by density functional theory calculations, we access the intrinsic picosecond response of peptide assemblies beyond the limits of conventional microscopy and ensemble averaging. Heterochiral (D-L) nanotubes display sharp resonances, whereas homochiral (L-L) nanotubes exhibit a comparatively featureless sub-THz profile, revealing a clear chirality-dependent contrast. Mode assignments show that the heterochiral features are dominated by localized torsional and bending motions of phenyl rings relative to the peptide backbone. Overall, chirality-dependent THz fingerprints emerge as sensitive descriptors of peptide nanotube architecture, while low-frequency nanospectroscopy provides a route to interrogate biomaterial interfaces under biologically relevant conditions. Authors Rajat Kumar, Erica Scarel, Andrea Perucchi, Lisa Vaccari, Prasanta Kumar Datta, Francesco D’amico, Paola Di Pietro, Silvia Marchesan, Federica Piccirilli Journal ChemrXiv Publication date 11/03/2026 Consult the publication

Open Lab
01/03/2026

Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models

Abstract Large language models (LLMs) are increasingly consulted for historical information by citizens, journalists, and institutions, raising concerns about their tendency to reproduce or amplify historical revisionism: the distortion, omission, or reframing of established facts. We introduce HistoricalMisinfo, a curated dataset of  contested events from  countries, each paired with factual and revisionist narratives. To approximate real-world dissemination, we design  prompt scenarios per event, capturing diverse ways historical content is elicited and framed. Using this benchmark, we evaluate multiple medium-sized LLMs and find systematic vulnerabilities: the prevalence of revisionist outputs varies across models, countries, and prompt types. HistoricalMisinfo provides a practical foundation for auditing the reliability of generative systems and for developing safeguards against the spread of revisionist narratives. Authors Francesco Ortu, Joeun Yook, Punya Syon Pandey, Keenan Samway, Bernhard Schölkopf, Alberto Cazzaniga, Rada Mihalcea, Zhijing Jin Journal Workshop AI4Peace @ International Conference of Learning Representations (ICLR) 2026 Publication Date 01/03/2026 Consult the publication

03/12/2025

Density-Informed VAE (DiVAE): Reliable Log-Prior Probability via Density Alignment Regularization

Abstract We introduce Density-Informed VAE (DiVAE), a lightweight, data-driven regularizer that aligns the VAE log-prior probability logpZ(z) with a log-density estimated from data. Standard VAEs match latents to a simple prior, overlooking density structure in the data-space. DiVAE encourages the encoder to allocate posterior mass in proportion to data-space density and, when the prior is learnable, nudges the prior toward high-density regions. This is realized by adding a robust, precision-weighted penalty to the ELBO, incurring negligible computational overhead. On synthetic datasets, DiVAE (i) improves distributional alignment of latent log-densities to its ground truth counterpart, (ii) improves prior coverage, and (iii) yields better OOD uncertainty calibration. On MNIST, DiVAE improves alignment of the prior with external estimates of the density, providing better interpretability, and improves OOD detection for learnable priors. Authors Michele Alessi, Alessio Ansuini, Alex Rodriguez Journal Principles of Generative Modeling (PriGM) Workshop @EurIPS2025 Publication Date 03/12/2025 Consult the publication

28/11/2025

Are LLMs Good Safety Agents or a Propaganda Engine?

Abstract Large Language Models (LLMs) are trained to refuse to respond to harmful content. However, systematic analyses of whether this behavior is truly a reflection of its safety policies or an indication of political censorship, that is practiced globally by countries, is lacking. Differentiating between safety influenced refusals or politically motivated censorship is hard and unclear. For this purpose we introduce PSP, a dataset built specifically to probe the refusal behaviors in LLMs from an explicitly political context. PSP is built by formatting existing censored content from two data sources, openly available on the internet: sensitive prompts in China generalized to multiple countries, and tweets that have been censored in various countries. We study: 1) impact of political sensitivity in seven LLMs through data-driven (making PSP implicit) and representation-level approaches (erasing the concept of politics); and, 2) vulnerability of models on PSP through prompt injection attacks (PIAs). Associating censorship with refusals on content with masked implicit intent, we find that most LLMs perform some form of censorship. We conclude with summarizing major attributes that can cause a shift in refusal distributions across models and contexts of different countries. Authors Neemesh Yadav, Francesco Ortu, Jiarui Liu, Joeun Yook, Bernhard Schölkopf, Rada Mihalcea, Alberto Cazzaniga, Zhijing Jin Journal Preprint Publication Date 28/11/2025 Consult the publication

20/11/2025

Harvesting energy from turbulent winds with reinforcement learning

Abstract Airborne Wind Energy (AWE) is an emerging technology designed to harness the power of high-altitude winds, offering a solution to several limitations of conventional wind turbines. AWE is based on flying devices (usually gliders or kites) that, tethered to a ground station and driven by the wind, convert its mechanical energy into electrical energy by means of a generator. Such systems are usually controlled by manoeuvering the kite so as to follow a predefined path prescribed by optimal control techniques, such as model-predictive control. These methods are strongly dependent on the specific model at use and difficult to generalize, especially in unpredictable conditions such as the turbulent atmospheric boundary layer. Our aim is to explore the possibility of replacing these techniques with an approach based on Reinforcement Learning (RL). Unlike traditional methods, RL does not require a predefined model, making it robust to variability and uncertainty. Our experimental results in complex simulated environments demonstrate that AWE agents trained with RL can effectively extract energy from turbulent flows, relying on minimal local information about the kite orientation and speed relative to the wind. Authors L. Basile, M. G. Berni, A. Celani Journal Europhysics Letters Publication Date 20/11/2025 Consult the publication

15/11/2025

Seeking Cost-Optimal Infrastructure Size for Distributed Filesystems: A Ceph Case Study

Abstract Distributed Filesystems (DFS) are a crucial component of modern computing environments, and their performance is critical to the success of all the facilities that rely on them. However, predicting the DFS I/O performance solely based on the storage system hardware is not trivial. In this paper, we address this challenge by presenting an empirical method that tries to quantitatively assess how hardware configuration choices influence the performance of a DFS using Ceph as a case study. We investigate the influence of three hardware parameters—number of CPU cores, amount of RAM, and disk bandwidth. To control these variables, we relied on the Linux hotplug interface and Cgroups, avoiding additional software overhead. Our results reveal that for the analyzed workloads, decreasing hardware resources does not always yield proportional performance losses. This method offers practical insights for designing cost-effective distributed storage systems, remaining general enough to be applied to other filesystems. Authors Niccolo Tosato, Isac Pasianotto, Ruggero Lot, Stefano Cozzini Journal Proceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis Publication Date 15/11/2025 Consult the publication

14/11/2025

Frequency maps reveal the correlation between Adversarial Attacks and Implicit Bias

Abstract Despite their impressive performance in classification tasks, neural networks are known to be vulnerable to adversarial attacks, subtle perturbations of the input data designed to deceive the model. In this work, we investigate the correlation between these perturbations and the implicit bias of neural networks trained with gradient-based algorithms. To this end, we analyse a representation of the network’s implicit bias through the lens of the Fourier transform. Specifically, we identify unique fingerprints of implicit bias and adversarial attacks by calculating the minimal, essential frequencies needed for accurate classification of each image, as well as the frequencies that drive misclassification in its adversarially perturbed counterpart. This approach enables us to uncover and analyse the correlation between these essential frequencies, providing a precise map of how the network’s biases align or contrast with the frequency components exploited by adversarial attacks. To this end, among other methods, we use a newly introduced technique capable of detecting nonlinear correlations between high-dimensional datasets. Our results provide empirical evidence that the network bias in Fourier space and the target frequencies of adversarial attacks are highly correlated and suggest new potential strategies for adversarial defence. Code is available at https://github.com/lorenzobasile/ImplicitBiasAdversarial Authors Lorenzo Basile, Nikos Karantzas, Alberto D’Onofrio, Luca Manzoni, Luca Bortolussi, Alex Rodriguez, Fabio Anselmi Journal 2025 International Joint Conference on Neural Networks (IJCNN) Publication Date 14/11/2025 Consult the publication

05/11/2025

A comprehensive framework for solution space exploration in community detection

Abstract: Community detection algorithms are essential tools for understanding complex networks, yet their results often vary between runs and are affected by node input order and the presence of outliers, undermining reproducibility and interpretation. This paper addresses these issues by introducing a framework for systematic exploration of the solution space, obtained through repeated runs of a given algorithm with permuted node orders. A Bayesian model assesses convergence, estimates solution probabilities, and provides a defensible stopping rule that balances accuracy and computational cost. Building on this process, we propose a taxonomy of solution spaces that offers clear diagnostics of partition reliability across algorithms and a shared vocabulary for interpretation. Applied to a real-world network, the approach shows that different algorithms produce various types of solution space, highlighting the importance of systematic exploration of the solutions before drawing scientific conclusions. Authors Fabio Morea, Domenico de Stefano Journal Scientific Reports Publication date 31/10/2025 Consult the pubblication  

03/10/2025

Ultrafast Intermolecular Dynamics of Nanoconfined Water in Swollen Lipid Cubic Mesophases

Abstract Understanding the structure and dynamics of the hydrogen-bond network ofwater in topologically distinct swollen lipidic mesophases, is fundamental fortheir application in biomedical, pharmaceutical, and food science fields. Here,a positive and non-linear correlation between water hydrogen-bond dynamicsand interfacial water population is uncovered in inverse bicontinuous swollenmesophases across an extended temperature range (298–340 K). Particularly,small-angle X-ray scattering determines the mesophase’s structural features,uncovering a temperature-driven re-entrant phenomenon (reappearance) ofPn̄ 3m phase upon heating. This topologically rich environment, however, hasno detectable impact on the temperature dependence of the intermolecularmodes of water, as revealed by terahertz absorption spectroscopy. Specifically,these modes show distinct dynamics: the stretching mode exhibits longerlifetimes than the libration mode, yet with a higher temperature-dependence,with approximately two-fold lower Arrhenius activation energies. In contrast,both stretching and libration modes exhibit a monotonic decrease in lifetimewith increasing temperature, due to the increasing disruption of thehydrogen-bond network. Atomistic molecular dynamics simulations enablethe quantification of interfacial water population, which shows a positivecorrelation with intermolecular lifetimes in a nonlinear manner, revealing anon-additive coupling between interfacial water population and waterhydrogen-bond network dynamics within these systems. Authors Eva Zunzunegui-Bru, Serena Rosa Alfarano, Patrick Züblin, Laura Baraldi,Hendrik Vondracek, Federica Piccirilli, Lisa Vaccari, and Raffaele Mezzenga Journal Small Publication date 03/10/2025 Consult the publication

Open Lab
18/09/2025

The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models

Abstract Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-language models (VLMs) handle image-understanding tasks, focusing on how visual information is processed and transferred to the textual domain. We compare native multimodal VLMs, models trained from scratch on multimodal data to generate both text and images, and non-native multimodal VLMs, models adapted from pre-trained large language models or capable of generating only text, highlighting key differences in information flow. We find that in native multimodal VLMs, image and text embeddings are more separated within the residual stream. Moreover, VLMs differ in how visual information reaches text: non-native multimodal VLMs exhibit a distributed communication pattern, where information is exchanged through multiple image tokens, whereas models trained natively for joint image and text generation tend to rely on a single post-image token that acts as a narrow gate for visual information. We show that ablating this single token significantly deteriorates image-understanding performance, whereas targeted, token-level interventions reliably steer image semantics and downstream text with fine-grained control. Authors Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon, Lucrezia Valeriani, Lorenzo Basile, Alessio Ansuini, Diego Doimo, Alberto Cazzaniga Journal Neurips 2025 Conference Publication Date 18/09/2025 Consult the publication

18/09/2025

Head Pursuit: Probing Attention Specialization in Multimodal Transformers

Abstract Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we study how individual attention heads in text-generative models specialize in specific semantic or visual attributes. Building on an established interpretability method, we reinterpret the practice of probing intermediate activations with the final decoding layer through the lens of signal processing. This lets us analyze multiple samples in a principled way and rank attention heads based on their relevance to target concepts. Our results show consistent patterns of specialization at the head level across both unimodal and multimodal transformers. Remarkably, we find that editing as few as 1% of the heads, selected using our method, can reliably suppress or enhance targeted concepts in the model output. We validate our approach on language tasks such as question answering and toxicity mitigation, as well as vision-language tasks including image classification and captioning. Our findings highlight an interpretable and controllable structure within attention layers, offering simple tools for understanding and editing large-scale generative models. Authors Lorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello, Alberto Cazzaniga Journal NeurIPS 2025 Conference (Spotlight) Publication Date 18/09/2025 Consult the publication