~/matteospanio

publications

23 entries · 6 years · 42 recorded citations across 17 tracked papers

what the papers talk about

audio audio — 724 occurrences across 20 of 21 papers tape tape — 400 occurrences across 9 of 21 papers model model — 533 occurrences across 19 of 21 papers music music — 483 occurrences across 19 of 21 papers irregularity irregularity — 268 occurrences across 8 of 21 papers preservation preservation — 308 occurrences across 11 of 21 papers taste taste — 220 occurrences across 8 of 21 papers video video — 246 occurrences across 13 of 21 papers signal signal — 234 occurrences across 17 of 21 papers musical musical — 173 occurrences across 11 of 21 papers recording recording — 156 occurrences across 9 of 21 papers restoration restoration — 124 occurrences across 8 of 21 papers image image — 162 occurrences across 14 of 21 papers cultural heritage cultural heritage — 53 occurrences across 9 of 21 papers mark mark — 112 occurrences across 8 of 21 papers filter filter — 104 occurrences across 7 of 21 papers generative model generative model — 40 occurrences across 6 of 21 papers speed speed — 135 occurrences across 12 of 21 papers food food — 93 occurrences across 6 of 21 papers artificial intelligence artificial intelligence — 57 occurrences across 14 of 21 papers analysis analysis — 159 occurrences across 19 of 21 papers recipe recipe — 72 occurrences across 4 of 21 papers repository repository — 95 occurrences across 7 of 21 papers deep learning deep learning — 39 occurrences across 7 of 21 papers learning learning — 126 occurrences across 13 of 21 papers text text — 126 occurrences across 13 of 21 papers language language — 135 occurrences across 15 of 21 papers interface interface — 110 occurrences across 10 of 21 papers generation generation — 124 occurrences across 13 of 21 papers neural networks neural networks — 45 occurrences across 10 of 21 papers MPAI MPAI — 96 occurrences across 8 of 21 papers audio document audio document — 40 occurrences across 8 of 21 papers detection detection — 107 occurrences across 10 of 21 papers score score — 130 occurrences across 15 of 21 papers processing processing — 111 occurrences across 11 of 21 papers audio recordings audio recordings — 39 occurrences across 8 of 21 papers emotional emotional — 66 occurrences across 4 of 21 papers digitization digitization — 93 occurrences across 8 of 21 papers multimodal multimodal — 93 occurrences across 8 of 21 papers library library — 93 occurrences across 8 of 21 papers file file — 126 occurrences across 15 of 21 papers symbolic music symbolic music — 27 occurrences across 4 of 21 papers generative generative — 123 occurrences across 15 of 21 papers task task — 123 occurrences across 15 of 21 papers module module — 104 occurrences across 11 of 21 papers corpus corpus — 88 occurrences across 8 of 21 papers project project — 101 occurrences across 11 of 21 papers modality modality — 74 occurrences across 6 of 21 papers human human — 104 occurrences across 12 of 21 papers audio analyzer audio analyzer — 25 occurrences across 4 of 21 papers playback playback — 99 occurrences across 11 of 21 papers training training — 119 occurrences across 16 of 21 papers environment environment — 89 occurrences across 9 of 21 papers synthesis synthesis — 59 occurrences across 4 of 21 papers frame frame — 93 occurrences across 10 of 21 papers tape irregularity tape irregularity — 30 occurrences across 6 of 21 papers composer composer — 87 occurrences across 9 of 21 papers participant participant — 71 occurrences across 6 of 21 papers audio file audio file — 38 occurrences across 10 of 21 papers evaluation evaluation — 128 occurrences across 20 of 21 papers ARP ARP — 80 occurrences across 8 of 21 papers digitized digitized — 80 occurrences across 8 of 21 papers algorithm algorithm — 97 occurrences across 12 of 21 papers graph graph — 61 occurrences across 5 of 21 papers style style — 85 occurrences across 10 of 21 papers audio tape audio tape — 27 occurrences across 6 of 21 papers application application — 112 occurrences across 18 of 21 papers machine learning machine learning — 29 occurrences across 7 of 21 papers analyzer analyzer — 58 occurrences across 5 of 21 papers analog audio analog audio — 24 occurrences across 5 of 21 papers sweet sweet — 63 occurrences across 6 of 21 papers archive archive — 76 occurrences across 9 of 21 papers implementation implementation — 84 occurrences across 11 of 21 papers sensory sensory — 62 occurrences across 6 of 21 papers computer computer — 80 occurrences across 10 of 21 papers code code — 91 occurrences across 13 of 21 papers access copy access copy — 23 occurrences across 5 of 21 papers irregularity classifier irregularity classifier — 23 occurrences across 5 of 21 papers music generation music generation — 23 occurrences across 5 of 21 papers playback speed playback speed — 25 occurrences across 6 of 21 papers audio audio — 724 occurrences across 20 of 21 papers tape tape — 400 occurrences across 9 of 21 papers model model — 533 occurrences across 19 of 21 papers music music — 483 occurrences across 19 of 21 papers irregularity irregularity — 268 occurrences across 8 of 21 papers preservation preservation — 308 occurrences across 11 of 21 papers taste taste — 220 occurrences across 8 of 21 papers video video — 246 occurrences across 13 of 21 papers signal signal — 234 occurrences across 17 of 21 papers musical musical — 173 occurrences across 11 of 21 papers recording recording — 156 occurrences across 9 of 21 papers restoration restoration — 124 occurrences across 8 of 21 papers image image — 162 occurrences across 14 of 21 papers cultural heritage cultural heritage — 53 occurrences across 9 of 21 papers mark mark — 112 occurrences across 8 of 21 papers filter filter — 104 occurrences across 7 of 21 papers generative model generative model — 40 occurrences across 6 of 21 papers speed speed — 135 occurrences across 12 of 21 papers food food — 93 occurrences across 6 of 21 papers artificial intelligence artificial intelligence — 57 occurrences across 14 of 21 papers analysis analysis — 159 occurrences across 19 of 21 papers recipe recipe — 72 occurrences across 4 of 21 papers repository repository — 95 occurrences across 7 of 21 papers deep learning deep learning — 39 occurrences across 7 of 21 papers learning learning — 126 occurrences across 13 of 21 papers text text — 126 occurrences across 13 of 21 papers language language — 135 occurrences across 15 of 21 papers interface interface — 110 occurrences across 10 of 21 papers generation generation — 124 occurrences across 13 of 21 papers neural networks neural networks — 45 occurrences across 10 of 21 papers MPAI MPAI — 96 occurrences across 8 of 21 papers audio document audio document — 40 occurrences across 8 of 21 papers detection detection — 107 occurrences across 10 of 21 papers score score — 130 occurrences across 15 of 21 papers
The 8034 terms most characteristic of 21 published papers.
how this is made
Terms are counted in the LaTeX sources, after stripping markup, maths and citations. Size follows how heavily a term is used by the papers that use it, damped by how many of them do — which ranks subject matter above the vocabulary every paper shares. Hover a term for its counts.

2026

2026IS2

A Quantized Native Runtime for On-Device Semantic Audio Generation

Matteo Spanio, Antonio Rodà

Proceedings of the IEEE International Symposium on the Internet of Sounds (IS2)

abstract

Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present <i>aria</i>, a dependency-free native runtime that runs the complete text-to-music pipeline of Stable Audio 3 (SA3) on ordinary GPUs, CPU-only machines, and a Raspberry Pi 5, with no Python or deep-learning framework underneath. Our main contribution is a study of <i>quantization</i>: running the model at lower numerical precision to fit tight memory budgets, saving memory in place rather than adding to it. Because the runtime owns every internal tensor, it also exposes activation steering, a low-cost way to steer what the model generates. We judge the quality cost with three independent measures of the output (prompt adherence, overall audio quality, taste preservation), each compared against the ordinary variation between random seeds. Eight-bit precision shows no measurable quality loss on any measure while sharply cutting memory, and it is the fastest mode on the GPU; four-bit adds a small, bounded cost but shrinks the footprint enough to run the 1.2-billion-parameter model on an 8 GB Pi. Against the official implementation, aria matches or exceeds generation speed and starts about seven times faster. A case study of the steering interface generates music carrying taste associations (<i>sonic seasoning</i>), with genuine but bounded control for a subset of attributes. These results make a compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings. The aria runtime is released at https://github.com/matteospanio/aria

bibtex
@inproceedings{spanio2026quantizednativeruntimeondevice,
      abbr = {IS2},
      code = {https://github.com/matteospanio/aria},
      pdf = {https://arxiv.org/pdf/2607.08526.pdf},
      title={A Quantized Native Runtime for On-Device Semantic Audio Generation},
      booktitle={Proceedings of the IEEE International Symposium on the Internet of Sounds (IS2)},
      author={Matteo Spanio and Antonio Rodà},
      year={2026},
      eprint={2607.08526},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      keywords={preprint},
      url={https://arxiv.org/abs/2607.08526},
      abstract={Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present \textit{aria}, a dependency-free native runtime that runs the complete text-to-music pipeline of Stable Audio~3 (SA3) on ordinary GPUs, CPU-only machines, and a Raspberry~Pi~5, with no Python or deep-learning framework underneath. Our main contribution is a study of \emph{quantization}: running the model at lower numerical precision to fit tight memory budgets, saving memory in place rather than adding to it. Because the runtime owns every internal tensor, it also exposes activation steering, a low-cost way to steer what the model generates. We judge the quality cost with three independent measures of the output (prompt adherence, overall audio quality, taste preservation), each compared against the ordinary variation between random seeds. Eight-bit precision shows no measurable quality loss on any measure while sharply cutting memory, and it is the fastest mode on the GPU; four-bit adds a small, bounded cost but shrinks the footprint enough to run the $1.2$-billion-parameter model on an $8$\,GB Pi. Against the official implementation, aria matches or exceeds generation speed and starts about seven times faster. A case study of the steering interface generates music carrying taste associations (\emph{sonic seasoning}), with genuine but bounded control for a subset of attributes. These results make a compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings. The aria runtime is released at https://github.com/matteospanio/aria}
}

%% SOTTOMESSO alla conferenza IEEE International Symposium on the Internet of Sounds (IS2) 2026

2026SMC1 citation

BMdataset: A Musicologically Curated LilyPond Dataset

Matteo Spanio, Ilay Guler, Antonio Rodà

Proceedings of the 23rd Sound and Music Computing Conference (SMC)

abstract

Symbolic music research has relied predominantly on MIDI-based datasets; text-based engraving formats such as LilyPond remain unexplored for music understanding. We present BMdataset, a musicologically curated dataset of 391 LilyPond scores (2,645 movements) transcribed by experts directly from original Baroque manuscripts, with metadata covering composer, musical form, instrumentation, and sectional attributes. Building on this resource, we introduce <b>LilyBERT</b>\footnote<a href="https://huggingface.co/csc-unipd/lilybert">https://huggingface.co/csc-unipd/lilybert</a>, a CodeBERT-based encoder adapted to symbolic music through vocabulary extension with 115 LilyPond-specific tokens and masked language model pre-training. Linear probing on the out-of-domain Mutopia corpus shows that, despite its modest size (∼90M tokens), fine-tuning on BMdataset alone outperforms in terms of accuracy continuous pre-training on the full PDMX corpus (∼15B tokens) for both composer and style classification, demonstrating that small, expertly curated datasets can be more effective than large, noisy corpora for music understanding. Combining broad pre-training with domain-specific fine-tuning yields the best results overall (84.3% composer accuracy), confirming that the two data regimes are complementary. We release the dataset, tokenizer, and model to establish a baseline for representation learning on LilyPond.

bibtex
@inproceedings{spanio2026bmdataset,
  language={en},
  abbr = {SMC},
  code = {https://github.com/CSCPadova/lilybert},
  arxiv = {2604.10628},
  pdf = {https://arxiv.org/pdf/2604.10628.pdf},
  bibtex_show={true},
  title={BMdataset: A Musicologically Curated LilyPond Dataset},
  author={Spanio, Matteo and Guler, Ilay and Rod{\`a}, Antonio},
  booktitle = {Proceedings of the 23rd Sound and Music Computing Conference (SMC)},
  year={2026},
  google_scholar_id={TQgYirikUcIC},
  abstract = {Symbolic music research has relied predominantly on MIDI-based datasets; text-based engraving formats such as LilyPond remain unexplored for music understanding. We present BMdataset, a musicologically curated dataset of 391 LilyPond scores (2,645 movements) transcribed by experts directly from original Baroque manuscripts, with metadata covering composer, musical form, instrumentation, and sectional attributes. Building on this resource, we introduce \textbf{LilyBERT}\footnote{\url{https://huggingface.co/csc-unipd/lilybert}}, a CodeBERT-based encoder adapted to symbolic music through vocabulary extension with 115 LilyPond-specific tokens and masked language model pre-training. Linear probing on the out-of-domain Mutopia corpus shows that, despite its modest size (${\sim}$90M tokens), fine-tuning on BMdataset alone outperforms in terms of accuracy continuous pre-training on the full PDMX corpus (${\sim}$15B tokens) for both composer and style classification, demonstrating that small, expertly curated datasets can be more effective than large, noisy corpora for music understanding. Combining broad pre-training with domain-specific fine-tuning yields the best results overall (84.3\% composer accuracy), confirming that the two data regimes are complementary. We release the dataset, tokenizer, and model to establish a baseline for representation learning on LilyPond.}
}

2026Ital-IA

Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding

Matteo Spanio, Mohammad Torabi, Andrea Poltronieri, Antonio Rodà

Proceedings of the Italian Conference on Artificial Intelligence (Ital-IA 2026)

abstract

Symbolic music evaluation for large language models remains fragmented across representations, datasets, and metrics. We introduce LilyBench, a LilyPond-based benchmark that jointly evaluates symbolic music generation and music understanding on the same family of open-weight LLMs. The benchmark includes a 200-prompt generation suite and ten understanding tasks adapted from ABC-Eval, covering syntax, metadata prediction, structural sequencing, and music recognition. Generation quality is evaluated using compile rate, MusPy descriptor distributions via Jensen–Shannon similarity, and LilyBERT-based Fréchet Music Distance (FMD). Experiments on four open-weight models show that executable LilyPond generation is achievable in zero-shot settings, while structural understanding tasks remain challenging despite strong performance on composer and genre recognition. Our experiments also reveal systematic disagreements between descriptor-based and embedding-based metrics, suggesting that symbolic music evaluation benefits from metric triangulation rather than single-score ranking. We release the benchmark, prompt bank, and evaluation code to support future research in symbolic music generation and understanding at https://github.com/CSCPadova/lilybench.

bibtex
@inproceedings{spanio2026lilypond,
  language={en},
  abbr = {Ital-IA},
  code = {https://github.com/CSCPadova/lilybench},
  pdf = {https://arxiv.org/pdf/2606.08722.pdf},
  bibtex_show={true},
  arxiv = {2606.08722},
  author    = {Matteo Spanio and Mohammad Torabi and Andrea Poltronieri and Antonio Rod{\`a}},
  title     = {Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding},
  booktitle = {Proceedings of the Italian Conference on Artificial Intelligence (Ital-IA 2026)},
  year      = {2026},
  publisher = {CEUR Workshop Proceedings},
  address   = {Italy},
  abstract = {Symbolic music evaluation for large language models remains fragmented across representations, datasets, and metrics. We introduce LilyBench, a LilyPond-based benchmark that jointly evaluates symbolic music generation and music understanding on the same family of open-weight LLMs. The benchmark includes a 200-prompt generation suite and ten understanding tasks adapted from ABC-Eval, covering syntax, metadata prediction, structural sequencing, and music recognition. Generation quality is evaluated using compile rate, MusPy descriptor distributions via Jensen–Shannon similarity, and LilyBERT-based Fréchet Music Distance (FMD). Experiments on four open-weight models show that executable LilyPond generation is achievable in zero-shot settings, while structural understanding tasks remain challenging despite strong performance on composer and genre recognition. Our experiments also reveal systematic disagreements between descriptor-based and embedding-based metrics, suggesting that symbolic music evaluation benefits from metric triangulation rather than single-score ranking. We release the benchmark, prompt bank, and evaluation code to support future research in symbolic music generation and understanding at https://github.com/CSCPadova/lilybench.},
}

2026

Cross-cultural evaluation of taste-sound correspondences in AI-generated music

Matteo Spanio, Massimiliano Zampini, Luisa Torri, Riccardo Migliavada, Bruno Mesz, Masaki Ohno,

+2 moreWada Yuji, Antonio Rodà

PLOS ONE

abstract

Sonic seasoning research has shown that listeners attribute systematic gustatory and emotional meaning to sound, and text-to-music generative artificial intelligence (AI) has recently been used to render gustatory prompts as musical stimuli. Whether the taste–sound correspondences acquired by such models hold beyond the cultural context in which they were validated remains untested. We extended a single-country Italian study to a three-country online experiment conducted in Argentina, Italy, and Japan (N=361), cohorts selected to span three continents and three distinct culinary and musical traditions. Participants first indicated their preference between base and fine-tuned MusicGen excerpts generated from four taste prompts (sweet, sour, bitter, salty), and then rated fine-tuned excerpts on twelve taste, emotion, and thermal descriptors. Preference for the fine-tuned model was confirmed in Argentina and Italy but not in Japan, and the salty prompt yielded the weakest correspondence in all three cohorts. Ratings differed substantially between countries, yet the main effect of country was no longer detectable once ratings had been standardized within participant, whereas the interactions characterizing the mapping of prompts onto descriptors remained essentially unchanged. Much of the apparent cross-cultural divergence is therefore attributable to differences in scale use; a structural component nevertheless persists. Excerpts occupying the same acoustic region, characterized by high spectral roughness and sensory dissonance, were predominantly labelled sour in Italy and Japan but bitter in Argentina, and exploratory factor analysis indicated that the twelve descriptors were organized along different latent dimensions in each cohort. These results indicate that cross-cultural variation in AI-mediated sonic seasoning operates at two levels: the overall level at which taste is attributed to a given stimulus, and the relational structure of those attributions. Evaluations of generative music systems across populations should accordingly distinguish response-style bias from genuine perceptual reorganization.

bibtex
@article{spanio2026,
  abbr = {PLOS ONE},
  title = {Cross-cultural evaluation of taste-sound correspondences in AI-generated music},
  author = {Spanio, Matteo and Zampini, Massimiliano and Torri, Luisa and Migliavada, Riccardo and Mesz, Bruno and Ohno, Masaki and Yuji, Wada and Rodà, Antonio},
  journal = {PLOS ONE},
  year = {2026},
  note = {Under review},
  abstract = {Sonic seasoning research has shown that listeners attribute systematic gustatory and emotional meaning to sound, and text-to-music generative artificial intelligence (AI) has recently been used to render gustatory prompts as musical stimuli. Whether the taste--sound correspondences acquired by such models hold beyond the cultural context in which they were validated remains untested. We extended a single-country Italian study to a three-country online experiment conducted in Argentina, Italy, and Japan ($N = 361$), cohorts selected to span three continents and three distinct culinary and musical traditions. Participants first indicated their preference between base and fine-tuned MusicGen excerpts generated from four taste prompts (sweet, sour, bitter, salty), and then rated fine-tuned excerpts on twelve taste, emotion, and thermal descriptors. Preference for the fine-tuned model was confirmed in Argentina and Italy but not in Japan, and the salty prompt yielded the weakest correspondence in all three cohorts. Ratings differed substantially between countries, yet the main effect of country was no longer detectable once ratings had been standardized within participant, whereas the interactions characterizing the mapping of prompts onto descriptors remained essentially unchanged. Much of the apparent cross-cultural divergence is therefore attributable to differences in scale use; a structural component nevertheless persists. Excerpts occupying the same acoustic region, characterized by high spectral roughness and sensory dissonance, were predominantly labelled sour in Italy and Japan but bitter in Argentina, and exploratory factor analysis indicated that the twelve descriptors were organized along different latent dimensions in each cohort. These results indicate that cross-cultural variation in AI-mediated sonic seasoning operates at two levels: the overall level at which taste is attributed to a given stimulus, and the relational structure of those attributions. Evaluations of generative music systems across populations should accordingly distinguish response-style bias from genuine perceptual reorganization.}
}

2026ISMIR

Does Notation Matter? MusicXML Tokenisation for Symbolic Music Generation and Understanding

Andrea Poltronieri, Matteo Spanio, Simone Chieppa, Xavier Serra, Martín Rocamora

Proceedings of the 27th International Society for Music Information Retrieval Conference (ISMIR)

abstract

Symbolic music representations encode musical content at varying levels of completeness, from performance-oriented formats such as MIDI to fully notated scores such as MusicXML. Yet the impact of this representational richness on deep learning models remains poorly understood. While tokenisation has been extensively studied for MIDI, score-based formats have received considerably less attention, and MusicXML in particular remains largely unexplored despite being the de-facto standard for score interchange. In this paper, we introduce the first general-purpose tokenisation framework for MusicXML, and conduct a systematic empirical study to determine whether notational richness confers a measurable advantage for deep learning on symbolic music. We train 26 autoregressive and masked language models, comparing MusicXML tokenisation against MIDI and ABC-based representations, probing the effect of different tokenisation strategies within MusicXML, and testing whether music-grammar-aware tokenisation outperforms treating scores as plain text. We evaluate music understanding tasks using linear probing across composer classification, difficulty estimation, and accompaniment suggestion, and generative tasks using objective statistical metrics and a human pairwise preference study. Our results show that MusicXML-tokenised models outperform MIDI and text-based counterparts on understanding and generative tasks.

bibtex
@inproceedings{poltronieri2026notation,
  abbr = {ISMIR},
  author    = {Poltronieri, Andrea and Spanio, Matteo and Chieppa, Simone and Serra, Xavier and Rocamora, Mart\'{i}n},
  title     = {Does Notation Matter? {MusicXML} Tokenisation for Symbolic Music Generation and Understanding},
  booktitle = {Proceedings of the 27th International Society for Music Information Retrieval Conference (ISMIR)},
  year      = {2026},
  address   = {Abu Dhabi, United Arab Emirates},
  abstract = {Symbolic music representations encode musical content at varying levels of completeness, from performance-oriented formats such as MIDI to fully notated scores such as MusicXML. Yet the impact of this representational richness on deep learning models remains poorly understood. While tokenisation has been extensively studied for MIDI, score-based formats have received considerably less attention, and MusicXML in particular remains largely unexplored despite being the de-facto standard for score interchange. In this paper, we introduce the first general-purpose tokenisation framework for MusicXML, and conduct a systematic empirical study to determine whether notational richness confers a measurable advantage for deep learning on symbolic music. We train 26 autoregressive and masked language models, comparing MusicXML tokenisation against MIDI and ABC-based representations, probing the effect of different tokenisation strategies within MusicXML, and testing whether music-grammar-aware tokenisation outperforms treating scores as plain text. We evaluate music understanding tasks using linear probing across composer classification, difficulty estimation, and accompaniment suggestion, and generative tasks using objective statistical metrics and a human pairwise preference study. Our results show that MusicXML-tokenised models outperform MIDI and text-based counterparts on understanding and generative tasks.}
}

%% SOTTOMESSO alla conferenza IEEE International Symposium on the Internet of Sounds (IS2) 2026, preprint arxiv

2026IS2

Fluxion: A Differentiable DSP Runtime from Training GPU to Real-Time Edge

Matteo Spanio, Antonio Rodà

Proceedings of the IEEE International Symposium on the Internet of Sounds (IS2)

abstract

Audio DSP tools still split into two camps. Native real-time engines meet audio deadlines but do not interoperate with modern machine-learning workflows. Tensor frameworks integrate well with machine learning, but they treat audio as offline batched tensors and provide no execution guarantees; runtime layers built on top of them inherit the host's limits in scheduling, GPU validation, and deployment footprint. This paper presents Fluxion, a Rust audio DSP runtime that reverses that trade. Fluxion <i>owns</i> the DSP-specific pieces that determine correctness and deployability: closed-form filter design, forward kernels, analytic adjoints for exact gradients, and stability certificates. It <i>rents</i> the general-purpose infrastructure that existing ML systems already provide: the autodiff graph, portable GPU code generation, and array interchange. One graph algebra lowers to a batched differentiable training engine and an allocation-free real-time engine. The two are connected by a certified freeze boundary that turns a trained graph into a versioned artifact that an edge device can re-certify before playback. In a co-run evaluation against six framework-hosted and native baselines, Fluxion leads the exact-recurrence systems on resident-GPU throughput and on most CPU workloads. Its real-time path carries a structural no-allocation guarantee rather than only a measured latency tail. The full train-certify-deploy loop also runs end to end on a Raspberry Pi 5, where a commercial restaurant smart-speaker workload is reproduced in real time on device. Fluxion is open source at https://github.com/matteospanio/fluxion

bibtex
@inproceedings{spanio2026fluxion,
      title={Fluxion: A Differentiable DSP Runtime from Training GPU to Real-Time Edge},
      code = {https://github.com/matteospanio/fluxion},
      author={Matteo Spanio and Antonio Rodà},
      year={2026},
      booktitle={Proceedings of the IEEE International Symposium on the Internet of Sounds (IS2)},
      abbr = {IS2},
      abstract = {Audio DSP tools still split into two camps. Native real-time engines meet audio deadlines but do not interoperate with modern machine-learning workflows. Tensor frameworks integrate well with machine learning, but they treat audio as offline batched tensors and provide no execution guarantees; runtime layers built on top of them inherit the host's limits in scheduling, GPU validation, and deployment footprint. This paper presents Fluxion, a Rust audio DSP runtime that reverses that trade. Fluxion \emph{owns} the DSP-specific pieces that determine correctness and deployability: closed-form filter design, forward kernels, analytic adjoints for exact gradients, and stability certificates. It \emph{rents} the general-purpose infrastructure that existing ML systems already provide: the autodiff graph, portable GPU code generation, and array interchange. One graph algebra lowers to a batched differentiable training engine and an allocation-free real-time engine. The two are connected by a certified freeze boundary that turns a trained graph into a versioned artifact that an edge device can re-certify before playback. In a co-run evaluation against six framework-hosted and native baselines, Fluxion leads the exact-recurrence systems on resident-GPU throughput and on most CPU workloads. Its real-time path carries a structural no-allocation guarantee rather than only a measured latency tail. The full train-certify-deploy loop also runs end to end on a Raspberry~Pi~5, where a commercial restaurant smart-speaker workload is reproduced in real time on device. Fluxion is open source at https://github.com/matteospanio/fluxion}
}

%% Sottomesso ad Artificial Intelligence Review

2026

Food and Flavor Representation Learning: A Systematic Mapping Study and Research Agenda

Matteo Spanio, Antonio Rodà

Artificial Intelligence Review

abstract

Machine learning for food, flavor, and odor increasingly replaces hand-engineered features with learned vector representations of molecules, ingredients, recipes, images, and sensor signals. We present a systematic mapping study of this shift: we search six bibliographic databases for work published between 2015 and 2026, identify 1648 records, screen 588 unique studies after deduplication, and retain 110 studies whose principal contribution is a learned representation, each coded for representation type, architecture, and data. The retained studies fall into three branches: molecular representations of taste and odor compounds (53 studies), food, recipe, and ingredient representations (54), and sensory or biosignal representations (3). Four patterns recur across the corpus: architecture choice tracks data modality rather than publication year; cross-modal recipe and image learning concentrates on a single dominant benchmark; audio, video, and physiological signals are barely represented; and transfer from foundation models stays fragmented across molecular, recipe, and sensory settings. We translate these observations into a research agenda and release the review protocol, the per-paper classification, and a lightweight quality rubric as an auditable companion package. The synthesis offers a descriptive map of the field rather than a formal study-level quality appraisal.

bibtex
@article{spanio2026review,
  abbr = {AIR},
  author = {Spanio, Matteo and Rodà, Antonio},
  title = {Food and Flavor Representation Learning: A Systematic Mapping Study and Research Agenda},
  journal = {Artificial Intelligence Review},
  year = {2026},
  note = {Under review},
  abstract = {Machine learning for food, flavor, and odor increasingly replaces hand-engineered features with learned vector representations of molecules, ingredients, recipes, images, and sensor signals.  We present a systematic mapping study of this shift: we search six bibliographic databases for work published between 2015 and 2026, identify 1648 records, screen 588 unique studies after deduplication, and retain 110 studies whose principal contribution is a learned representation, each coded for representation type, architecture, and data.  The retained studies fall into three branches: molecular representations of taste and odor compounds (53 studies), food, recipe, and ingredient representations (54), and sensory or biosignal representations (3).  Four patterns recur across the corpus: architecture choice tracks data modality rather than publication year; cross-modal recipe and image learning concentrates on a single dominant benchmark; audio, video, and physiological signals are barely represented; and transfer from foundation models stays fragmented across molecular, recipe, and sensory settings.  We translate these observations into a research agenda and release the review protocol, the per-paper classification, and a lightweight quality rubric as an auditable companion package. The synthesis offers a descriptive map of the field rather than a formal study-level quality appraisal.}
}

%% Sottomesso a PLOS ONE

2026arXiv1 citation

Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences

Matteo Spanio, Valentina Frezzato, Antonio Rodà

arXiv preprint arXiv:2604.10632

abstract

Music and food traditions are both intangible cultural heritage, and the links between them, how a sound can make a taste seem sweeter or more bitter, are increasingly used in museum, exhibition and gastronomic-tourism settings. Modelling those links computationally runs into a data bottleneck familiar across cultural heritage computing: expert annotation is slow and costly, so the annotated collections that result are small. The usual remedy is to enlarge a collection automatically, labelling it with a model trained on the small annotated one. Such <i>synthetic</i> labels are rarely checked, either against the original annotations or against people. We provide both checks. Experiment 1 asks whether the audio–flavour patterns found in an experimental soundtrack collection (257 tracks annotated by listeners) survive when the collection is scaled to ∼49,300 30-second segments from the Free Music Archive (FMA) labelled by a fine-tuned Audio Spectrogram Transformer. Experiment 2 asks whether flavour profiles computed from food chemistry, for 20 dishes drawn largely from Italian culinary tradition, match what listeners actually hear (49 participants, online). Feature–flavour patterns carry over for every taste dimension (ρ=0.38–0.72, all p<0.001), and sweetness still carries over when every spectral feature is removed, so the agreement is not an artefact of how the labelling model represents audio. Listener ratings match the computed profiles far beyond chance (permutation p<0.001; Mantel r=0.45; Procrustes m²=0.49), and the result holds when participants reporting hearing or taste impairments are excluded. We release the harmonized datasets and all code.

bibtex
@article{spanio2026multimodal,
  language={en},
  code = {https://github.com/CSCPadova/music-flavor-analysis.git},
  arxiv = {2604.10632},
  pdf = {https://arxiv.org/pdf/2604.10632.pdf},
  bibtex_show={true},
  title={Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences},
  author={Spanio, Matteo and Frezzato, Valentina and Rod{\`a}, Antonio},
  journal={arXiv preprint arXiv:2604.10632},
  year={2026},
  google_scholar_id={HDshCWvjkbEC},
  abstract={Music and food traditions are both intangible cultural heritage, and the links between them, how a sound can make a taste seem sweeter or more bitter, are increasingly used in museum, exhibition and gastronomic-tourism settings. Modelling those links computationally runs into a data bottleneck familiar across cultural heritage computing: expert annotation is slow and costly, so the annotated collections that result are small. The usual remedy is to enlarge a collection automatically, labelling it with a model trained on the small annotated one. Such \emph{synthetic} labels are rarely checked, either against the original annotations or against people. We provide both checks. Experiment~1 asks whether the audio--flavour patterns found in an experimental soundtrack collection (257~tracks annotated by listeners) survive when the collection is scaled to $\sim$49,300 30-second segments from the Free Music Archive (FMA) labelled by a fine-tuned Audio Spectrogram Transformer. Experiment~2 asks whether flavour profiles computed from food chemistry, for 20 dishes drawn largely from Italian culinary tradition, match what listeners actually hear (49~participants, online). Feature--flavour patterns carry over for every taste dimension ($\rho=0.38$--$0.72$, all $p<0.001$), and sweetness still carries over when every spectral feature is removed, so the agreement is not an artefact of how the labelling model represents audio. Listener ratings match the computed profiles far beyond chance (permutation $p<0.001$; Mantel $r=0.45$; Procrustes $m^{2}=0.49$), and the result holds when participants reporting hearing or taste impairments are excluded. We release the harmonized datasets and all code.}
}

2026Ital-IA

Preserving Memory, Expanding Creativity: Human-Centered AI Trajectory in Engineering and Music Research at the CSC of Padua University

Sergio Canazza, Alessandro Fiordelmondo, Sara Giuriati, Cristina Paulon, Niccolò Pretto, Antonio Rodà,

+3 moreAlessandro Russo, Matteo Spanio, Anna Zuccante

Proceedings of the Italian Conference on Artificial Intelligence (Ital-IA 2026)

abstract

Artificial intelligence is reshaping musical creation, analysis, performance, and preservation, but its relationship with music has a history that predates today’s generative models. This paper presents the Centro di Sonologia Computazionale (CSC) of the University of Padua as a long-standing example of human-centered AI research in music. Since 1979, the CSC has promoted interdisciplinary collaboration among artists, engineers, scholars, archivists, and cognitive scientists. The paper argues that AI should not be understood as a replacement for human creativity, but as a means to expand musical intelligence, enable co-creative practices, and protect cultural heritage. It focuses on three areas: Disklavier-based composer–AI collaboration, multimodal generative systems informed by psychophysics, and AI-supported preservation of computer music and analog audio archives with commercial applications. Together, these examples show how AI can support musical innovation while preserving interpretability, historical awareness, and human aesthetic agency.

bibtex
@inproceedings{canazza2026preserving,
  language={en},
  abbr = {Ital-IA},
  bibtex_show={true},
  author    = {Sergio Canazza and Alessandro Fiordelmondo and
               Sara Giuriati and Cristina Paulon and
               Niccol{\`o} Pretto and Antonio Rod{\`a} and
               Alessandro Russo and Matteo Spanio and
               Anna Zuccante},
  title     = {Preserving Memory, Expanding Creativity: Human-Centered AI Trajectory in Engineering and Music Research at the CSC of Padua University},
  booktitle = {Proceedings of the Italian Conference on Artificial Intelligence (Ital-IA 2026)},
  year      = {2026},
  publisher = {CEUR Workshop Proceedings},
  abstract = {Artificial intelligence is reshaping musical creation, analysis, performance, and preservation, but its relationship with music has a history that predates today’s generative models. This paper presents the Centro di Sonologia Computazionale (CSC) of the University of Padua as a long-standing example of human-centered AI research in music. Since 1979, the CSC has promoted interdisciplinary collaboration among artists, engineers, scholars, archivists, and cognitive scientists. The paper argues that AI should not be understood as a replacement for human creativity, but as a means to expand musical intelligence, enable co-creative practices, and protect cultural heritage. It focuses on three areas: Disklavier-based composer--AI collaboration, multimodal generative systems informed by psychophysics, and AI-supported preservation of computer music and analog audio archives with commercial applications. Together, these examples show how AI can support musical innovation while preserving interpretability, historical awareness, and human aesthetic agency.},
}

2026CBMI1 citation

Taste-aware music retrieval from audio embeddings

Matteo Spanio, Antonio Rodá

Proceedings of the International Conference on Content-Based Multimedia Indexing (CBMI)

abstract

Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs.\ 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single \encoderVGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs.\ 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound–taste correspondences.

bibtex
@inproceedings{spanio2026cbmi,
    abbr = {CBMI},
    author = {Matteo Spanio and Antonio Rod{\'a}},
    title = {Taste-aware music retrieval from audio embeddings},
    year = {2026},
    booktitle = {Proceedings of the  International Conference on Content-Based Multimedia Indexing (CBMI)},
    google_scholar_id={j3f4tGmQtD8C},
    abstract = {Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of $0.134$; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE $0.13$ vs.\ $0.28$), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline ($0.219$). On absolute error the encoders are statistically flat, with a single \encoder{VGGish} matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson $r$ $0.724$ vs.\ $0.666$). Operationalised as a content-based retrieval index, the predicted taste space ranks a $309$-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound--taste correspondences.}
}

2026AVI-CH

Toward an Advanced Web Interface for Analyzing and Annotating Digitized Audio Tape Recordings

Niccolò Pretto, Andrea Franceschini, Alessandro Russo, Matteo Spanio, Sergio Canazza

Proceedings of the AVI-CH 2026 Workshop on Advanced Visual Interfaces for Cultural Heritage

abstract

The access and study of digitized historical audio recordings require more than simple sound playback, as these materials embody complex physical and contextual information. Previous studies have explored the virtualization of tape recorder functionalities; however, such approaches alone are insufficient to fully support scholarly analysis and restoration processes. Recent developments in standards, such as IEEE/MPAI Standard 3302-2024 on Audio Recording Preservation, have emphasized the potential of integrating automated multimodal analysis to support restoration and access, including valuable contributions from artificial intelligence techniques. The main contribution of this work is a standalone web interface for analyzing and annotating digitized historical audio recordings, integrating synchronized audio and video of the tape during the digitization process to provide additional insights, such as annotations and the presence of splices. The proposed interface represents an initial step toward the integration of MPAI-based technologies for restoration. In this context, a computer vision approach is also introduced and evaluated, capable of identifying key components such as the playback head, magnetic tape, and pinch roller, which are essential for supporting future extensions toward (semi-)automated restoration techniques.

bibtex
@inproceedings{pretto2026advanced,
  language={en},
  abbr = {AVI-CH},
  bibtex_show={true},
  author    = {Niccol{\`o} Pretto and Andrea Franceschini and
               Alessandro Russo and Matteo Spanio and Sergio Canazza},
  title     = {Toward an Advanced Web Interface for Analyzing and Annotating Digitized Audio Tape Recordings},
  booktitle = {Proceedings of the AVI-CH 2026 Workshop on Advanced Visual Interfaces for Cultural Heritage},
  year      = {2026},
  publisher = {CEUR Workshop Proceedings},
  abstract = {The access and study of digitized historical audio recordings require more than simple sound playback, as these materials embody complex physical and contextual information. Previous studies have explored the virtualization of tape recorder functionalities; however, such approaches alone are insufficient to fully support scholarly analysis and restoration processes. Recent developments in standards, such as IEEE/MPAI Standard 3302-2024 on Audio Recording Preservation, have emphasized the potential of integrating automated multimodal analysis to support restoration and access, including valuable contributions from artificial intelligence techniques. The main contribution of this work is a standalone web interface for analyzing and annotating digitized historical audio recordings, integrating synchronized audio and video of the tape during the digitization process to provide additional insights, such as annotations and the presence of splices. The proposed interface represents an initial step toward the integration of MPAI-based technologies for restoration. In this context, a computer vision approach is also introduced and evaluated, capable of identifying key components such as the playback head, magnetic tape, and pinch roller, which are essential for supporting future extensions toward (semi-)automated restoration techniques.}
}

2025

2025Frontiers CSselected7 citations

A multimodal symphony: integrating taste and sound through generative AI

Matteo Spanio, Massimiliano Zampini, Antonio Rodà, Franco Pierucci

Frontiers in Computer Science

abstract

In recent decades, neuroscientific and psychological research has identified direct relationships between taste and auditory perception. This article explores multimodal generative models capable of converting taste information into music, building on this foundational research. We provide a brief review of the state of the art in this field, highlighting key findings and methodologies. We present an experiment in which a fine-tuned version of a generative music model (MusicGEN) is used to generate music based on detailed taste descriptions provided for each musical piece. The results are promising: according to the participants' evaluations (n = 111), the fine-tuned model produces music that more coherently reflects the input taste descriptions compared to the non-fine-tuned model. This study represents a significant step toward understanding and developing embodied interactions between AI, sound, and taste, opening new possibilities in the field of generative AI.

bibtex
@article{spanio_frontiers_2025,
  altmetric = {true},
  dimensions = {true},
  abbr = {Frontiers in CS},
  selected ={true},
  preview={frontiers2025.webp},
  bibtex_show={true},
  language={en},
  author={Spanio, Matteo and Zampini, Massimiliano and Rodà, Antonio and Pierucci, Franco},
  title={A multimodal symphony: integrating taste and sound through generative AI},
  journal={Frontiers in Computer Science},
  volume={Volume 7 - 2025},
  year={2025},
  url={https://www.frontiersin.org/journals/computer-science/articles/10.3389/fcomp.2025.1575741},
  pdf={frontiers2025.pdf},
  doi={10.3389/fcomp.2025.1575741},
  issn={2624-9898},
  abstract={In recent decades, neuroscientific and psychological research has identified direct relationships between taste and auditory perception. This article explores multimodal generative models capable of converting taste information into music, building on this foundational research. We provide a brief review of the state of the art in this field, highlighting key findings and methodologies. We present an experiment in which a fine-tuned version of a generative music model (MusicGEN) is used to generate music based on detailed taste descriptions provided for each musical piece. The results are promising: according to the participants' evaluations (n = 111), the fine-tuned model produces music that more coherently reflects the input taste descriptions compared to the non-fine-tuned model. This study represents a significant step toward understanding and developing embodied interactions between AI, sound, and taste, opening new possibilities in the field of generative AI.},
  google_scholar_id={WF5omc3nYNoC}
}

2025DAFx

TorchFX: A Modern Approach to Audio DSP with PyTorch and GPU Acceleration

Matteo Spanio, Antonio Rodà

Proceedings of the 28th International Conference on Digital Audio Effects (DAFx)

abstract

The increasing complexity and real-time processing demands of audio signals require optimized algorithms that utilize the computational power of Graphics Processing Units (GPUs). Existing Digital Signal Processing (DSP) libraries often do not provide the necessary efficiency and flexibility, particularly for integrating with Artificial Intelligence (AI) models. In response, we introduce TorchFX: a GPU-accelerated Python library for DSP, engineered to facilitate sophisticated audio signal processing. Built on the PyTorch framework, TorchFX offers an Object-Oriented interface similar to torchaudio but enhances functionality with a novel pipe operator for intuitive filter chaining. The library provides a comprehensive suite of Finite Impulse Response (FIR) and Infinite Impulse Response (IIR) filters, with a focus on multichannel audio, thereby facilitating the integration of DSP and AI-based approaches. Our benchmarking results demonstrate significant efficiency gains over traditional libraries like SciPy, particularly in multichannel contexts. While there are current limitations in GPU compatibility, ongoing developments promise broader support and real-time processing capabilities. TorchFX aims to become a useful tool for the community, contributing to innovation in GPU-accelerated DSP. TorchFX is publicly available on GitHub at https://github.com/matteospanio/torchfx. © 2025 Matteo Spanio et al.

bibtex
@inproceedings{spanio2025torchfx,
  language={en},
  poster = {dafx_poster_2025.pdf},
  abbr = {DAFX},
	iris={true},
  bibtex_show={true},
  altmetric = {true},
  website = {https://matteospanio.github.io/torchfx/},
  arxiv = {2504.08624},
	author = {Spanio, Matteo and Rodà, Antonio},
	title = {{TorchFX}: A Modern Approach to Audio {DSP} with {PyTorch} and {GPU} Acceleration},
	year = {2025},
	booktitle = {Proceedings of the 28th International Conference on Digital Audio Effects (DAFx)},
	pages = {390--395},
	url = {https://www.scopus.com/inward/record.uri?eid=2-s2.0-105028935688&partnerID=40&md5=552e54afc1a074cbd1b7e8ed4ad1c010},
  pdf={https://dafx25.dii.univpm.it/wp-content/uploads/2025/07/DAFx25_paper_65.pdf},
	abstract = {The increasing complexity and real-time processing demands of audio signals require optimized algorithms that utilize the computational power of Graphics Processing Units (GPUs). Existing Digital Signal Processing (DSP) libraries often do not provide the necessary efficiency and flexibility, particularly for integrating with Artificial Intelligence (AI) models. In response, we introduce TorchFX: a GPU-accelerated Python library for DSP, engineered to facilitate sophisticated audio signal processing. Built on the PyTorch framework, TorchFX offers an Object-Oriented interface similar to torchaudio but enhances functionality with a novel pipe operator for intuitive filter chaining. The library provides a comprehensive suite of Finite Impulse Response (FIR) and Infinite Impulse Response (IIR) filters, with a focus on multichannel audio, thereby facilitating the integration of DSP and AI-based approaches. Our benchmarking results demonstrate significant efficiency gains over traditional libraries like SciPy, particularly in multichannel contexts. While there are current limitations in GPU compatibility, ongoing developments promise broader support and real-time processing capabilities. TorchFX aims to become a useful tool for the community, contributing to innovation in GPU-accelerated DSP. TorchFX is publicly available on GitHub at https://github.com/matteospanio/torchfx. © 2025 Matteo Spanio et al.},
}

2025NIME3 citations

Towards a Repository Template for Music Technology Research

Alessandro Fiordelmondo, Matteo Spanio, Patricia Cadavid, Xinran Chen, Sergio Canazza, Raul Masu

Proceedings of the International Conference on New Interfaces for Musical Expression

abstract

Documenting and sharing research output is essential to construct the critical discourse on new music technology. Documentation feeds the knowledge and the values with which to evaluate and discuss current achievements and musical creations as well as to plan for the future. Besides publishing our research in conferences and journals, sharing research materials and outcomes like software, hardware, instruments, and datasets is important. This allows others to use the latest technology and improve it. For this purpose, the repository is increasingly commonly used by researchers and artists to store and share their works. However, creating repositories does not follow a clear and organised structure like the one we find, for example, in papers. The heterogeneity of repositories makes it hard to use both practically and for analysis. Although the variety and differences of research products in the field of new musical technologies are obvious, we believe that defining repositories with common guidelines could significantly improve the critical discourse in this area. This issue has been discussed at the NIME conference through workshops and papers. In this article, we want to continue this discussion and propose a flexible repository template to organise and present research materials and outcomes in the field of musical technologies research. The article provides a short and focused review of how repositories are currently used at the NIME conference, with special attention to the platforms used. Based on this study, we introduce a repository template that will be applied to case studies. We hope this proposal will encourage further discussion and advancement on this issue and, at the same time, support and facilitate the creation of new repositories.

bibtex
@inproceedings{fiordelmondo_nime_2025,
  language={en},
  abbr = {NIME},
  code = {https://github.com/CSCPadova/MTR-template},
  bibtex_show = {true},
  author       = {Fiordelmondo, Alessandro and
                  Spanio, Matteo and
                  Cadavid, Patricia and
                  Chen, Xinran and
                  Canazza, Sergio and
                  Masu, Raul},
  title        = {Towards a Repository Template for Music Technology
                   Research
                  },
  booktitle    = {Proceedings of the International Conference on New
                   Interfaces for Musical Expression
                  },
  year         = {2025},
  pages        = {556--562},
  publisher    = {Zenodo},
  venue        = {Canberra, Australia},
  doi          = {10.5281/zenodo.15698958},
  pdf          = {nime2025_81.pdf},
  url          = {https://doi.org/10.5281/zenodo.15698958},
  google_scholar_id = {LkGwnXOMwfcC},
  abstract = {Documenting and sharing research output is essential to construct the critical discourse on new music technology. Documentation feeds the knowledge and the values with which to evaluate and discuss current achievements and musical creations as well as to plan for the future. Besides publishing our research in conferences and journals, sharing research materials and outcomes like software, hardware, instruments, and datasets is important. This allows others to use the latest technology and improve it. For this purpose, the repository is increasingly commonly used by researchers and artists to store and share their works. However, creating repositories does not follow a clear and organised structure like the one we find, for example, in papers. The heterogeneity of repositories makes it hard to use both practically and for analysis. Although the variety and differences of research products in the field of new musical technologies are obvious, we believe that defining repositories with common guidelines could significantly improve the critical discourse in this area. This issue has been discussed at the NIME conference through workshops and papers. In this article, we want to continue this discussion and propose a flexible repository template to organise and present research materials and outcomes in the field of musical technologies research. The article provides a short and focused review of how repositories are currently used at the NIME conference, with special attention to the platforms used. Based on this study, we introduce a repository template that will be applied to case studies. We hope this proposal will encourage further discussion and advancement on this issue and, at the same time, support and facilitate the creation of new repositories.}
}

2024

2024AESselected1 citation

A novel derivative-based approach for the automatic detection of time-reversed audio in the MPAI/IEEE-CAE ARP international standard

Marina Bosi, Fabio Zanini, Matteo Spanio, Alessandro Russo, Canazza Sergio

Proceedings of the 157th Audio Engineering Society Convention (AES)

abstract

The Moving Picture, Audio and Data Coding by Artificial Intelligence (MPAI) Context-based Audio Enhancement (CAE) Audio Recording Preservation (ARP) standard provides the technical specifications for a comprehensive framework for digitizing and preserving analog audio, specifically focusing on documents recorded on open-reel tapes. This paper presents a novel envelope derivative-based method designed to be integrated into the ARP standard, for detecting reverse audio sections during the preservation process. The primary objective of this method is to automatically identify segments of audio recorded in reverse. Leveraging derivative-based signal processing algorithms, the system enhances its capability to detect and reverse such sections, thereby reducing errors during the preservation process. This feature not only aids in identifying and correcting errors but also enhances the efficiency of large-scale audio document archiving projects. The system's performance was evaluated using a diverse dataset that includes various musical genres and digitized tapes, demonstrating its strong potential and effectiveness across different types of audio content.

bibtex
@inproceedings{bosi2024a,
  language= {en},
  dimensions = {true},
  preview={aes2024.jpg},
  abbr = {AES},
  bibtex_show={true},
  selected = {true},
  code = {https://github.com/CSCPadova/reverse-detection},
  author={Bosi, Marina
    and Zanini, Fabio
    and Spanio, Matteo
    and Russo, Alessandro
    and Canazza Sergio},
  booktitle = {Proceedings of the 157th Audio Engineering Society Convention (AES)},
  title={A novel derivative-based approach for the automatic detection of time-reversed audio in the MPAI/IEEE-CAE ARP international standard},
  year={2024},
  number={10190},
  abstract={The Moving Picture, Audio and Data Coding by Artificial Intelligence (MPAI) Context-based Audio Enhancement (CAE) Audio Recording Preservation (ARP) standard provides the technical specifications for a comprehensive framework for digitizing and preserving analog audio, specifically focusing on documents recorded on open-reel tapes. This paper presents a novel envelope derivative-based method designed to be integrated into the ARP standard, for detecting reverse audio sections during the preservation process. The primary objective of this method is to automatically identify segments of audio recorded in reverse. Leveraging derivative-based signal processing algorithms, the system enhances its capability to detect and reverse such sections, thereby reducing errors during the preservation process. This feature not only aids in identifying and correcting errors but also enhances  the efficiency of large-scale audio document archiving projects. The system's performance was evaluated using a diverse dataset that includes various musical genres and digitized tapes, demonstrating its strong potential and effectiveness across different types of audio content.},
  url={https://aes2.org/publications/elibrary-page/?id=22693},
  google_scholar_id={qjMakFHDy7sC},
}

2024ICIAPselected7 citations

Enhancing Preservation and Restoration of Open Reel Audio Tapes Through Computer Vision

Alessandro Russo, Matteo Spanio, Sergio Canazza

Image Analysis and Processing - ICIAP 2023 Workshops

abstract

Analog audio documents inevitably face degradation over time, posing a challenge for preserving their audio content and ensuring the integrity of the recordings. Analog document preservation is one of the main research topics of interest of the Centro di Sonologia Computazionale (CSC) of the Department of Information Engineering of the University of Padua, which over the years developed and implemented a methodology for preservation that includes, among other things, the video recording of the digitization process of the open-reel tapes for documenting irregularities on the top of their surface. Together with the corpus of digitized high-quality audio recordings, this led to the creation of an internal archive of video documents. This paper presents a software application that leverages computer vision techniques to automatically detect Irregularities on open-reel audio tapes, analyzing the video documents produced during the digitization interventions. The software employs a frame-by-frame analysis to automatically identify and highlight points of interest that may indicate tape damages, splices, and other Irregularities. The software uses Generalized Hough Transform and SURF algorithms to locate regions of interest within the tape. The proposed software is also part of the MPAI/IEEE-CAE ARP standard developed by Audio Innova s.r.l., spin-off of the CSC, and it may offer a robust and efficient solution for analyzing open-reel audio tapes, supporting archivists and musicologists in their activities.

bibtex
@inproceedings{10.1007/978-3-031-51026-7_26,
  abbr = {ICIAP},
  dimensions = {true},
  selected = {true},
  preview = {scene_obj.png},
  bibtex_show={true},
  language= {en},
  doi={10.1007/978-3-031-51026-7_26},
  url={https://doi.org/10.1007/978-3-031-51026-7_26},
  author={Russo, Alessandro
    and Spanio, Matteo
    and Canazza, Sergio},
  editor={Foresti, Gian Luca
    and Fusiello, Andrea
    and Hancock, Edwin},
  title={Enhancing Preservation and Restoration of Open Reel Audio Tapes Through Computer Vision},
  booktitle={Image Analysis and Processing - ICIAP 2023 Workshops},
  year={2024},
  publisher={Springer Nature Switzerland},
  address={Cham},
  pages={297--308},
  abstract={Analog audio documents inevitably face degradation over time, posing a challenge for preserving their audio content and ensuring the integrity of the recordings. Analog document preservation is one of the main research topics of interest of the Centro di Sonologia Computazionale (CSC) of the Department of Information Engineering of the University of Padua, which over the years developed and implemented a methodology for preservation that includes, among other things, the video recording of the digitization process of the open-reel tapes for documenting irregularities on the top of their surface. Together with the corpus of digitized high-quality audio recordings, this led to the creation of an internal archive of video documents. This paper presents a software application that leverages computer vision techniques to automatically detect Irregularities on open-reel audio tapes, analyzing the video documents produced during the digitization interventions. The software employs a frame-by-frame analysis to automatically identify and highlight points of interest that may indicate tape damages, splices, and other Irregularities. The software uses Generalized Hough Transform and SURF algorithms to locate regions of interest within the tape. The proposed software is also part of the MPAI/IEEE-CAE ARP standard developed by Audio Innova s.r.l., spin-off of the CSC, and it may offer a robust and efficient solution for analyzing open-reel audio tapes, supporting archivists and musicologists in their activities.},
  isbn={978-3-031-51026-7},
  google_scholar_id={9yKSN-GCB0IC},
}

2024IAI4CH

Filming the sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms

Zafer Çınar, Alessandro Russo, Matteo Spanio, Niccolò Pretto, Sergio Canazza

Proceedings of the 3rd Workshop on Artificial Intelligence for Cultural Heritage (IAI4CH 2024) co-located with the 23rd International Conference of the Italian Association for Artificial Intelligence (AIxIA 2024)

abstract

The preservation of open-reel audio tapes is critical for maintaining valuable cultural and historical audio archives, yet current digitisation and analysis operations are often error-prone due to tape degradation and the long duration of the recordings. Considering the analog nature of this kind of recording, anomaly detection algorithms, applied to the video of the tape flowing on the playback head, can be used to detect errors and details with musicological value. This paper presents a new dataset of high-quality videos and a new algorithm for anomaly detection on audio tapes. Experimental results show notable improvements in detection performance, though false positives remain a challenge at higher speeds. Additionally, the new algorithm supports a wider range of playback speeds, improving its flexibility. This improvement is an important step towards a reliable implementation of the IEEE/MPAI CAE ARP standard (3302-2022).

bibtex
@inproceedings{Cinar2024,
  abbr = {IAI4CH},
  preview = {flowchart.png},
  bibtex_show={true},
  language = {en},
  author = {Çınar, Zafer and Russo, Alessandro and Spanio, Matteo and Pretto, Niccolò and Canazza, Sergio},
  title = {Filming the sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms},
  booktitle = {Proceedings of the 3rd Workshop on Artificial Intelligence for Cultural Heritage (IAI4CH 2024) co-located with the 23rd International Conference of the Italian Association for Artificial Intelligence (AIxIA 2024)},
  series ={{CEUR} Workshop Proceedings},
  year = {2024},
  abstract = {The preservation of open-reel audio tapes is critical for maintaining valuable cultural and historical audio archives, yet current digitisation and analysis operations are often error-prone due to tape degradation and the long duration of the recordings. Considering the analog nature of this kind of recording, anomaly detection algorithms, applied to the video of the tape flowing on the playback head, can be used to detect errors and details with musicological value. This paper presents a new dataset of high-quality videos and a new algorithm for anomaly detection on audio tapes. Experimental results show notable improvements in detection performance, though false positives remain a challenge at higher speeds. Additionally, the new algorithm supports a wider range of playback speeds, improving its flexibility. This improvement is an important step towards a reliable implementation of the IEEE/MPAI CAE ARP standard (3302-2022).},
  publisher = {http://CEUR-WS.org},
  url = {https://ceur-ws.org/Vol-3865/},
  pdf = {https://ceur-ws.org/Vol-3865/13_paper.pdf},
  html = {https://doi.org/10.5281/zenodo.14028923},
  doi = {10.5281/zenodo.14028923},
}

2024IEEE Access11 citations

From Tape to Code: An International AI-Based Standard for Audio Cultural Heritage Preservation - Don’t Play That Song for me (If it’s Not Preserved With ARP!)

Marina Bosi, Sergio Canazza, Niccolò Pretto, Alessandro Russo, Matteo Spanio

IEEE Access

abstract

This article describes a novel technology for preserving audio documents archived on open-reel magnetic tapes forming the core of the Audio Recording Preservation (ARP) international standard. ARP is part of the Moving Picture, Audio, and Data Coding by Artificial Intelligence (MPAI) Context-based Audio Enhancement (CAE) standard, adopted by the IEEE Standard Association as IEEE 3302-2022 in December 2022. Leveraging automated Artificial Intelligence (AI) tools, ARP analyzes and extracts relevant information from digitized audio and video files of the tape’s corresponding digital Preservation Copy. This process includes identifying speed variations and surface irregularities on the tape, automatically rectifying errors to generate a restored Access Copy. By utilizing the ARP standard, archives gain a potent tool for expediting and optimizing the description of the preservation conditions of the tape, as well as automatically correcting any errors that may have occurred during the digitization process. This technology offers an efficient solution for managing both small and large collections of digitized analog items, marking a substantial advancement in the preservation of audio documents.

bibtex
@article{ieeeaccess2024,
  language= {en},
  dimensions = {true},
  preview={access-gagraphic-3474529.jpg},
  abbr = {IEEEAccess},
  pdf={IEEE_Access_2024.pdf},
  bibtex_show={true},
  html={https://ieeexplore.ieee.org/document/10705421},
  author={Bosi, Marina and Canazza, Sergio and Pretto, Niccolò and Russo, Alessandro and Spanio, Matteo},
  journal={IEEE Access},
  title={From Tape to Code: An International AI-Based Standard for Audio Cultural Heritage Preservation - Don’t Play That Song for me (If it’s Not Preserved With ARP!)},
  year={2024},
  volume={12},
  number={},
  pages={152544-152558},
  keywords={Artificial intelligence;Music;Cultural differences;Guidelines;Production;Magnetic recording;Surface treatment;Magnetoacoustic effects;Audio recording;Accuracy;Document handling;Audio systems;Artificial intelligence;audio documents preservation;audio restoration;IEEE standard;musicological analysis;MPAI standard},
  abstract={This article describes a novel technology for preserving audio documents archived on open-reel magnetic tapes forming the core of the Audio Recording Preservation (ARP) international standard. ARP is part of the Moving Picture, Audio, and Data Coding by Artificial Intelligence (MPAI) Context-based Audio Enhancement (CAE) standard, adopted by the IEEE Standard Association as IEEE 3302-2022 in December 2022. Leveraging automated Artificial Intelligence (AI) tools, ARP analyzes and extracts relevant information from digitized audio and video files of the tape’s corresponding digital Preservation Copy. This process includes identifying speed variations and surface irregularities on the tape, automatically rectifying errors to generate a restored Access Copy. By utilizing the ARP standard, archives gain a potent tool for expediting and optimizing the description of the preservation conditions of the tape, as well as automatically correcting any errors that may have occurred during the digitization process. This technology offers an efficient solution for managing both small and large collections of digitized analog items, marking a substantial advancement in the preservation of audio documents.},
  doi={10.1109/ACCESS.2024.3474529},
  google_scholar_id = {2osOgNQ5qMEC},
}

2024AIxIA DC7 citations

Towards Emotionally Aware AI: Challenges and Opportunities in the Evolution of Multimodal Generative Models

Matteo Spanio

Proceedings of the AIxIA Doctoral Consortium 2024 co-located with the 23nd International Conference of the Italian Association for Artificial Intelligence (AIxIA 2024)

abstract

The evolution of generative models in artificial intelligence (AI) has significantly expanded the capacity of machines to process and generate complex multimodal data such as text, images, audio, and video. Despite these advancements, the integration of emotional awareness remains an underexplored dimension. This paper examines the state of the art in multimodal generative AI, with a focus on existing models developed by major technology companies. It then proposes an approach to incorporate emotional awareness into AI models, which would enhance human-machine interaction by improving the interpretability and explainability of AI-generated decisions. The paper also addresses the challenges associated with building emotion-aware models, including the need for comprehensive multimodal datasets and the computational complexity of incorporating less-explored sensory modalities like olfaction and gustation. Finally, potential solutions are discussed, including the normalization of existing research data and the application of transfer learning to reduce resource demands. These steps are essential for advancing the field and unlocking the potential of emotion-aware multimodal AI in applications such as healthcare, robotics, and virtual assistants.

bibtex
@inproceedings{Spanio2024,
  abbr = {AIxIADC},
  dimensions = {true},
  author = {Spanio, Matteo},
  poster={aixia_poster.pdf},
  bibtex_show={true},
  language = {en},
  title = {Towards Emotionally Aware AI: Challenges and Opportunities in the Evolution of Multimodal Generative Models},
  booktitle = {Proceedings of the AIxIA Doctoral Consortium 2024 co-located with the 23nd International Conference of the Italian Association for Artificial Intelligence (AIxIA 2024)},
  series = {{CEUR} Workshop Proceedings},
  year = {2024},
  abstract = { The evolution of generative models in artificial intelligence (AI) has significantly expanded the capacity of machines to process and generate complex multimodal data such as text, images, audio, and video. Despite these advancements, the integration of emotional awareness remains an underexplored dimension. This paper examines the state of the art in multimodal generative AI, with a focus on existing models developed by major technology companies. It then proposes an approach to incorporate emotional awareness into AI models, which would enhance human-machine interaction by improving the interpretability and explainability of AI-generated decisions. The paper also addresses the challenges associated with building emotion-aware models, including the need for comprehensive multimodal datasets and the computational complexity of incorporating less-explored sensory modalities like olfaction and gustation. Finally, potential solutions are discussed, including the normalization of existing research data and the application of transfer learning to reduce resource demands. These steps are essential for advancing the field and unlocking the potential of emotion-aware multimodal AI in applications such as healthcare, robotics, and virtual assistants.},
  publisher = {http://CEUR-WS.org},
  url = {https://ceur-ws.org/Vol-3914/},
  pdf = {https://ceur-ws.org/Vol-3914/short84.pdf},
  google_scholar_id = {W7OEmFMy1HYC}
}

2023

2023thesis2 citations

A study on Equalization Curve Detection in Audio Tape Digitization process using Artificial Intelligence

Matteo Spanio

abstract

In recent decades, archives have seen a rapid change in the media used to store sound information, and many of these media are rich in obsolete material that risks becoming unusable due to aging. Therefore, it is necessary to digitize sound documents in order to make them durable over time. However, during the digitization process, errors such as applying an incorrect equalization curve or playing back the tape at the wrong speed can lead to the acquisition of inauthentic material. This work focuses on studying the detection of possible errors due to incorrect equalization curve settings and tape playback speed during the transfer of material from analog to digital, verifying if and how it is possible to detect them using methods specific to Artificial Intelligence (clustering and classification). The results of this research demonstrate that these algorithms may offer good precision in detecting errors and have the potential to automate the verification process, ensuring the preservation of valid information for a longer period of time, but before they can be used in a real-world scenario, they must be further improved.

bibtex
@thesis{https://doi.org/10.13140/rg.2.2.36838.50247,
  bibtex_show={true},
  pdf = {A_study_on_Equalization_Curve_Detection.pdf},
  doi = {10.13140/RG.2.2.36838.50247},
  url = {https://rgdoi.net/10.13140/RG.2.2.36838.50247},
  html = {https://matteospanio.gitlab.io/mpai-audio-analyser/},
  author = {Spanio, Matteo},
  language = {en},
  title = {A study on Equalization Curve Detection in Audio Tape Digitization process using Artificial Intelligence},
  publisher = {Unpublished},
  year = {2023},
  google_scholar_id={d1gkVwhDpl0C},
  abstract = {In recent decades, archives have seen a rapid change in the media used to store sound information, and many of these media are rich in obsolete material that risks becoming unusable due to aging. Therefore, it is necessary to digitize sound documents in order to make them durable over time. However, during the digitization process, errors such as applying an incorrect equalization curve or playing back the tape at the wrong speed can lead to the acquisition of inauthentic material. This work focuses on studying the detection of possible errors due to incorrect equalization curve settings and tape playback speed during the transfer of material from analog to digital, verifying if and how it is possible to detect them using methods specific to Artificial Intelligence (clustering and classification). The results of this research demonstrate that these algorithms may offer good precision in detecting errors and have the potential to automate the verification process, ensuring the preservation of valid information for a longer period of time, but before they can be used in a real-world scenario, they must be further improved.},
}

2021

2021thesis1 citation

TUTTI QUANTI VOGLION FARE JAZZ - Contaminazioni Jazz nel repertorio clarinettistico del '900

Matteo Spanio

abstract

Nel complesso questo lavoro vuole essere un invito per i musicisti classici a riconsiderare la prevalenza di un genere su un altro: non solo Mozart, ma anche Benny Goodman. La musica tradizionale non è musica meno seria di quella definita “colta”, sono tutte espressioni di parti diverse della nostra società e possono esistere grandi artisti tanto in un genere quanto nell’altro. Non solo è necessario rivalutare i propri pregiudizi, ma è solo grazie a questo cambio di prospettiva che è possibile scoprire quanto swing sia presente anche nella musica del passato: solo chi ha fatto tesoro di esperienze classiche e jazz può utilizzare i modelli compositivi colti con il linguaggio appropriato, troppo spesso, in entrambe le direzioni, si è assistito a esperimenti tuttaltro che felici perché si partiva da presupposti incompleti. Ma in alcune occasioni si può apprezzare come la conoscenza di più generi ponga le basi per i migliori successi: si pensi al trombettista Wynton Marsalis che ha vinto contemporaneamente un Grammy Award per la musica classica e per il jazz nel 1994.

bibtex
@thesis{https://doi.org/10.13140/rg.2.2.17327.82088,
  bibtex_show={true},
  pdf = {TUTTI_QUANTI_VOGLION_FARE_JAZZ_Contamina.pdf},
  doi = {10.13140/RG.2.2.17327.82088},
  url = {https://rgdoi.net/10.13140/RG.2.2.17327.82088},
  author = {Spanio, Matteo},
  language = {it},
  title = {TUTTI QUANTI VOGLION FARE JAZZ - Contaminazioni Jazz nel repertorio clarinettistico del '900},
  type = {Master's Thesis},
  publisher = {Unpublished},
  google_scholar_id = {u-x6o8ySG0sC},
  year = {2021},
  abstract = {Nel complesso questo lavoro vuole essere un invito per i musicisti classici a riconsiderare la prevalenza di un genere su un altro: non solo Mozart, ma anche Benny Goodman. La musica tradizionale non è musica meno seria di quella definita “colta”, sono tutte espressioni di parti diverse della nostra società e possono esistere grandi artisti tanto in un genere quanto nell’altro. Non solo è necessario rivalutare i propri pregiudizi, ma è solo grazie a questo cambio di prospettiva che è possibile scoprire quanto swing sia presente anche nella musica del passato: solo chi ha fatto tesoro di esperienze classiche e jazz può utilizzare i modelli compositivi colti con il linguaggio appropriato, troppo spesso, in entrambe le direzioni, si è assistito a esperimenti tuttaltro che felici perché si partiva da presupposti incompleti. Ma in alcune occasioni si può apprezzare come la conoscenza di più generi ponga le basi per i migliori successi: si pensi al trombettista Wynton Marsalis che ha vinto contemporaneamente un Grammy Award per la musica classica e per il jazz nel 1994.}
}

2019

2019thesis

IL CLARINETTO ALL'OPERA

Matteo Spanio

abstract

Attraverso le ricerche riportate in questa tesi si è voluto approfondire nelle sue diverse sfaccettature il ruolo del clarinetto e dei clarinettisti nella composizione delle Fantasie su temi d’Opera, genere molto diffuso all’inizio dell’Ottocento. Si sono visti nel dettaglio il Potpourri n. 2 per clarinetto e orchestra su Là ci darem la mano di Franz Danzi, le Variazioni su Euer Liebreiz, eure Schönheit in Si♭ maggiore dall’Opera Alruna di Louis Spohr e la Fantasia da Concerto su motivi del Rigoletto di Luigi Bassi. Per ogni brano si sono analizzate le origini storiche ponendo grande attenzione al contatto e talvolta alla collaborazione avvenuta tra compositore e strumentista. Lo scopo di tale lavoro non è quello di riportare semplicemente nozioni di valore storico, ma di riuscire a fornire al lettore l’idea di una corretta interpretazione filologica grazie al supporto della ricostruzione della vita dei personaggi coinvolti nella storia di questi brani; tenendo conto del fatto che il paesaggio sonoro in cui viviamo è diverso da quello di duecento anni fa e che le caratteristiche acustiche del clarinetto hanno subito numerose modifiche nel corso del tempo. La tesi è articolata in quattro capitoli e un’appendice in cui si può trovare una breve storia del clarinetto. Ogni capitolo è strutturato in maniera indipendente e può essere letto separatamente dal resto della tesi. Per ogni capitolo viene fornita un’introduzione al contesto storico e geografico a cui si fa riferimento, una storia dell’autore e dell’esecutore del pezzo e una breve analisi del brano considerato.

bibtex
@thesis{https://doi.org/10.13140/rg.2.2.24143.56489,
  bibtex_show={true},
  pdf = {IL_CLARINETTO_ALLOPERA.pdf},
  abstract = {Attraverso le ricerche riportate in questa tesi si è voluto approfondire nelle sue diverse sfaccettature il ruolo del clarinetto e dei clarinettisti nella composizione delle Fantasie su temi d’Opera, genere molto diffuso all’inizio dell’Ottocento. Si sono visti nel dettaglio il Potpourri n. 2 per clarinetto e orchestra su Là ci darem la mano di Franz Danzi, le Variazioni su Euer Liebreiz, eure Schönheit in Si♭ maggiore dall’Opera Alruna di Louis Spohr e la Fantasia da Concerto su motivi del Rigoletto di Luigi Bassi. Per ogni brano si sono analizzate le origini storiche ponendo grande attenzione al contatto e talvolta alla collaborazione avvenuta tra compositore e strumentista. Lo scopo di tale lavoro non è quello di riportare semplicemente nozioni di valore storico, ma di riuscire a fornire al lettore l’idea di una corretta interpretazione filologica grazie al supporto della ricostruzione della vita dei personaggi coinvolti nella storia di questi brani; tenendo conto del fatto che il paesaggio sonoro in cui viviamo è diverso da quello di duecento anni fa e che le caratteristiche acustiche del clarinetto hanno subito numerose modifiche nel corso del tempo. La tesi è articolata in quattro capitoli e un’appendice in cui si può trovare una breve storia del clarinetto. Ogni capitolo è strutturato in maniera indipendente e può essere letto separatamente dal resto della tesi. Per ogni capitolo viene fornita un’introduzione al contesto storico e geografico a cui si fa riferimento, una storia dell’autore e dell’esecutore del pezzo e una breve analisi del brano considerato.},
  doi = {10.13140/RG.2.2.24143.56489},
  url = {https://rgdoi.net/10.13140/RG.2.2.24143.56489},
  author = {Spanio, Matteo},
  language = {it},
  title = {IL CLARINETTO ALL'OPERA},
  type = {Bachelor's Thesis},
  publisher = {Unpublished},
  year = {2019},
}

2019score

Variations on Theme from Alruna

Louis Spohr, Matteo Spanio

Edizioni Eufonia

bibtex
@misc{sphor:woo15,
  bibtex_show = {true},
  preview={variation-on-theme-from-arluna.jpg},
  author  = {Spohr, Louis and Spanio, Matteo},
  howpublished = {Edizioni Eufonia},
  title   = {Variations on Theme from Alruna},
  year         = {2019},
  url     = {https://www.edizionieufonia.it/homepage/variation-on-theme-from-arluna-for-clarinet-and-piano/},
  html     = {https://www.edizionieufonia.it/homepage/variation-on-theme-from-arluna-for-clarinet-and-piano/},
  note    = {Reduction for clarinet and piano by Matteo Spanio},
}