~/matteospanio

> whoami

Matteo Spanio

AI Research Engineer · University of Padua

Matteo Spanio

I’m an AI research engineer working on audio, music and multimodal generative models, and a PhD candidate in Brain, Mind and Computer Science at the University of Padua, based at the Centro di Sonologia Computazionale. In 2025 I spent six months as a visiting PhD student at the Music Technology Group of Universitat Pompeu Fabra in Barcelona.

Three threads run through the work: recovering audio cultural heritage from degraded magnetic tape, as part of the IEEE 3302-2022 standard I help develop at MPAI; teaching generative models the correspondences between taste and sound; and asking what language models actually understand about music notation.

Alongside it I maintain TorchFX, an audio DSP library that runs on the GPU, and I play clarinet with the Orchestra di Padova e del Veneto. The long version is in the CV.

research map

## corpus explorer

Everything I have written, placed by how similar the texts are. Click a node for its card, trace paths between works, search the corpus in plain words, turn on the sound to hear the map from wherever the camera stands — and drag the projection slider to see how much of this picture is real structure.

23 papers, 6 projects, 5 posts and 8 news, placed by how similar their text is.

## the research map

Everything I have written — 23 papers, 6 projects, 5 posts, 8 news — embedded with intfloat/multilingual-e5-base over 108 text chunks and projected to two dimensions with UMAP. Distance is textual similarity; the lines join each work's closest neighbours; colour is my declared research area, never a clustering of the embedding.

Here: drag to pan, scroll to zoom, click a node for its card — and from a card, trace the shortest chain of similarity links to any other work.

Sound (the speaker) plays the map rather than decorating it. The camera is the listener: pan and zoom move you across and above the constellation, so what you are looking at is what you hear nearby. A signal walks the similarity graph, one note a second, and each research area has its own instrument — tape flutter for the heritage work, a plucked string for the Baroque scores, a reed for musicology. Opening a work sounds it with its nearest neighbours; a traced path plays as a phrase.

Full screen turns the map into a small app: semantic search that runs entirely in your browser, and a slider across 12 alternative projections of the same embedding.

Read the geometry loosely. The areas separate only weakly here — silhouette 0.30 in this picture, 0.15 in the full embedding — because nearly everything is audio, music and machine learning (within-area similarity 0.86, between-area 0.82). The lines are computed in the embedding rather than in the picture, so who connects to whom means more than how far apart things sit.

AI for audio cultural heritage

Multimodal and crossmodal AI

Symbolic music and LilyPond

Audio DSP tooling

Musicology and performance

Software and research engineering

selected work

all 23 publications →

news

  • Ital-IA 2026
  • Lilybert is available on Hugging Face
  • DAFx 2025
  • New DL model released

all news →

recent writing

all posts →