> whoami
Matteo Spanio
AI Research Engineer · University of Padua

I’m an AI research engineer working on audio, music and multimodal generative models, and a PhD candidate in Brain, Mind and Computer Science at the University of Padua, based at the Centro di Sonologia Computazionale. In 2025 I spent six months as a visiting PhD student at the Music Technology Group of Universitat Pompeu Fabra in Barcelona.
Three threads run through the work: recovering audio cultural heritage from degraded magnetic tape, as part of the IEEE 3302-2022 standard I help develop at MPAI; teaching generative models the correspondences between taste and sound; and asking what language models actually understand about music notation.
Alongside it I maintain TorchFX, an audio DSP library that runs on the GPU, and I play clarinet with the Orchestra di Padova e del Veneto. The long version is in the CV.
research map
how this is made
Read the geometry loosely. The areas separate only weakly here — silhouette 0.30 — because nearly everything is audio, music and machine learning: within-area similarity 0.86, between-area 0.82. And the picture flatters them, since the same score in the full embedding is only 0.15. The lines are computed in that embedding rather than in the picture, so who sits next to whom means more than how far apart things look.
AI for audio cultural heritage
- Enhancing Preservation and Restoration of Open Reel Audio Tapes Through Computer Vision
- A study on Equalization Curve Detection in Audio Tape Digitization process using Artificial Intelligence
- A novel derivative-based approach for the automatic detection of time-reversed audio in the MPAI/IEEE-CAE ARP international standard
- From Tape to Code: An International AI-Based Standard for Audio Cultural Heritage Preservation - Don’t Play That Song for me (If it’s Not Preserved With ARP!)
- Filming the sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms
- Preserving Memory, Expanding Creativity: Human-Centered AI Trajectory in Engineering and Music Research at the CSC of Padua University
- Toward an Advanced Web Interface for Analyzing and Annotating Digitized Audio Tape Recordings
- Ital-IA 2026
- AES 2024
- ICIAP 2023
Multimodal and crossmodal AI
- Towards Emotionally Aware AI: Challenges and Opportunities in the Evolution of Multimodal Generative Models
- A multimodal symphony: integrating taste and sound through generative AI
- Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences
- Taste-aware music retrieval from audio embeddings
- Food and Flavor Representation Learning: A Systematic Mapping Study and Research Agenda
- Cross-cultural evaluation of taste-sound correspondences in AI-generated music
- tasty-musicgen
- AIxIA 2024
- New DL model released
Symbolic music and LilyPond
Audio DSP tooling
Musicology and performance
Software and research engineering
- Towards a Repository Template for Music Technology Research
- A Quantized Native Runtime for On-Device Semantic Audio Generation
- Fluxion: A Differentiable DSP Runtime from Training GPU to Real-Time Edge
- Principles of statistics
- Python's virtual environments
- Python's Static Typing Safari: In Search of Code Clarity
- ML Project template
- spam-analyzer
- ACDL
selected work
- 2025A multimodal symphony: integrating taste and sound through generative AI
Frontiers in Computer Science
- 2024A novel derivative-based approach for the automatic detection of time-reversed audio in the MPAI/IEEE-CAE ARP international standard
Proceedings of the 157th Audio Engineering Society Convention (AES)
- 2024Enhancing Preservation and Restoration of Open Reel Audio Tapes Through Computer Vision
Image Analysis and Processing - ICIAP 2023 Workshops
news
- Ital-IA 2026
- Lilybert is available on Hugging Face
- DAFx 2025
- New DL model released
recent writing
- Making scientific Python blazingly fast with PyTorch
Swapping NumPy and SciPy for PyTorch buys GPU acceleration, an object-oriented API and AI interoperability at once.
- Python's virtual environments
A tour of pip, venv, pipx, conda and poetry, and of what each one keeps out of your system Python.
- Python's Static Typing Safari: In Search of Code Clarity
Growing a one-line addition function into annotations, overloads and TypeVar generics, so type mismatches surface in the editor instead of at runtime.