A defining human trait is the ability to convey complex information through writing and drawing. This has enabled the creation of documents, the foundational medium for preserving and transmitting knowledge across time and space. This dissertation presents multimodal methods, datasets, benchmarks, and evaluation tools for understanding, generating, and editing Visually Rich Documents, spanning both cultural heritage sources and modern digital content. The term `Visually Rich Documents' refers to pages where text, figures, graphics, and layout jointly carry meaning. When dealing with historical sources, which are often centuries old, reliable analysis first requires reducing the complexity of the physical substrate to make downstream tasks tractable. Therefore, on the analysis side, we introduce a binarization and restoration model agnostic to language, script, support, and resolution, reducing bleed-through, noise, and material artifacts in 2D document images while preserving informative strokes. We further extend this spectral approach to 3D tomographic volumes of the Herculaneum papyri, demonstrating preliminary applicability to ink detection on carbonized ancient scrolls. As many Visually Rich Documents encode meaning through spatial organization, we develop a multi-page parsing system that extracts regions, relations, and visual semantics without relying on an external optical text recognition model, enabling the parsing of complex, multi-page layouts. To foster research in ultra-low-resource settings, we curate a dataset for the Coptic language and benchmark current methods. Having established methods for document understanding, we turn to generation, targeting the core constituents of Visually Rich Documents: styled text and figures. We analyze existing evaluation approaches for styled text generation and propose an evaluation score specifically sensitive to calligraphic and typographic style, addressing a systematic flaw in the use of generic image generation metrics for this task. We then introduce methods for controllable handwritten and typewritten styled text generation, conditioning on content, style, and geometric variation. For visual content with extreme aspect ratios, such as panoramas, banners, and wide illustrations, we adapt large-scale pretrained diffusion models in a training-free manner, coordinating multiple generation views to produce globally coherent outputs at arbitrary aspect ratios without introducing new trainable parameters. Additionally, we tackle illustrations, characterized by irregular shapes and semi-transparent elements, and propose a pipeline to create, from a textual prompt, transparency-aware (RGBA) assets ready for compositing in complex layouts. Finally, we address multi-instance infographic editing by introducing an attention-partitioning mechanism for flow matching models that disentangles concurrent edits across multiple localized regions, preventing attribute leakage and enabling single-pass editing. We accompany this contribution with the first large-scale benchmark for multi-instance infographic editing, comprising images with up to 285 simultaneous edit instances, an order of magnitude beyond prior benchmarks. With this thesis, we aim to help close the loop between analysis, generation, and editing, enabling new applications in archive digitization, dataset augmentation, and content creation for Visually Rich Documents, while showing that models, metrics, and benchmarks developed for one stage of the pipeline can directly support the others.

Un tratto distintivo dell'essere umano è la capacità di trasmettere informazioni complesse tramite scrittura e disegno. Questo ha reso possibile la creazione di documenti, il mezzo fondamentale per preservare e trasmettere la conoscenza attraverso il tempo e lo spazio. Questa tesi di dottorato presenta architetture multimodali di deep learning per la comprensione e la generazione di Documenti Visivamente Complessi, coprendo sia le fonti del patrimonio culturale che i contenuti digitali moderni. Il termine `Documenti Visivamente Complessi' si riferisce a pagine in cui testo, figure, grafica e layout contribuiscono congiuntamente al significato. Nel caso delle fonti storiche, che spesso hanno secoli di storia, un'analisi affidabile richiede di ridurre la complessità del substrato fisico per rendere trattabili i processi a valle. Sul fronte dell'analisi, introduciamo un modello di binarizzazione e restauro agnostico rispetto a lingua, sistema di scrittura, supporto e risoluzione, che riduce il passaggio d'inchiostro (bleed-through), il rumore e gli artefatti del materiale nelle immagini 2D, preservando i tratti informativi. Estendiamo inoltre questo approccio spettrale ai volumi tomografici 3D, dimostrando un'applicabilità preliminare al rilevamento dell'inchiostro su antichi rotoli carbonizzati. Poiché molti Documenti Visivamente Complessi codificano il significato tramite la disposizione spaziale, sviluppiamo un sistema di parsing multi-pagina che estrae regioni, relazioni e semantica visiva senza ricorrere a un modello esterno di OCR, consentendo il parsing di layout complessi e multi-pagina. Per promuovere la ricerca in contesti con risorse estremamente limitate (ultra-low-resource), curiamo un dataset per la lingua Copta ed effettuiamo il benchmarking dei metodi attuali. Una volta definiti i metodi per la comprensione dei documenti, passiamo alla generazione, concentrandoci sui componenti fondamentali dei Documenti Visivamente Complessi: il testo stilizzato e le figure. Analizziamo gli approcci di valutazione esistenti per la generazione di testo stilizzato e proponiamo uno score di valutazione sensibile allo stile calligrafico e tipografico, affrontando un limite sistematico nell'uso di metriche generiche per questo compito. Introduciamo poi metodi per la generazione controllabile di testo stilizzato manoscritto e dattiloscritto, condizionando su contenuto, stile e variazione geometrica. Per i contenuti visivi con proporzioni estreme, come panoramiche, banner e illustrazioni orizzontali, adattiamo modelli diffusivi pre-addestrati su larga scala in modo privo di addestramento aggiuntivo, coordinando molteplici viste di generazione per produrre output globalmente coerenti a proporzioni arbitrarie. Inoltre, affrontiamo le illustrazioni, caratterizzate da forme irregolari ed elementi semitrasparenti, e proponiamo un'architettura per creare, da un prompt testuale, asset con trasparenza (RGBA), pronti per il compositing in layout complessi. Infine, affrontiamo le infografiche, la forma visiva più sofisticata di informazione documentale, introducendo un meccanismo di partizionamento dell'attenzione per i modelli di flow matching che disaccoppia le modifiche concorrenti su più regioni localizzate, prevenendo la contaminazione degli attributi e abilitando l'editing multi-istanza in un unico passaggio. Presentiamo inoltre il primo benchmark su larga scala per questo compito, comprendente immagini con fino a 285 istanze da modificare simultaneamente. Con questa tesi, puntiamo a contribuire a chiudere il cerchio tra analisi e generazione, abilitando nuove applicazioni nell'arricchimento dei dataset (dataset augmentation), nella digitalizzazione degli archivi e nella creazione di contenuti per i Documenti Visivamente Complessi.

Architetture Multimodali per la Comprensione e la Generazione di Documenti Visivamente Complessi / Fabio Quattrini , 2026 Jul 20. 38. ciclo, Anno Accademico 2024/2025.

Architetture Multimodali per la Comprensione e la Generazione di Documenti Visivamente Complessi

QUATTRINI, FABIO
2026

Abstract

A defining human trait is the ability to convey complex information through writing and drawing. This has enabled the creation of documents, the foundational medium for preserving and transmitting knowledge across time and space. This dissertation presents multimodal methods, datasets, benchmarks, and evaluation tools for understanding, generating, and editing Visually Rich Documents, spanning both cultural heritage sources and modern digital content. The term `Visually Rich Documents' refers to pages where text, figures, graphics, and layout jointly carry meaning. When dealing with historical sources, which are often centuries old, reliable analysis first requires reducing the complexity of the physical substrate to make downstream tasks tractable. Therefore, on the analysis side, we introduce a binarization and restoration model agnostic to language, script, support, and resolution, reducing bleed-through, noise, and material artifacts in 2D document images while preserving informative strokes. We further extend this spectral approach to 3D tomographic volumes of the Herculaneum papyri, demonstrating preliminary applicability to ink detection on carbonized ancient scrolls. As many Visually Rich Documents encode meaning through spatial organization, we develop a multi-page parsing system that extracts regions, relations, and visual semantics without relying on an external optical text recognition model, enabling the parsing of complex, multi-page layouts. To foster research in ultra-low-resource settings, we curate a dataset for the Coptic language and benchmark current methods. Having established methods for document understanding, we turn to generation, targeting the core constituents of Visually Rich Documents: styled text and figures. We analyze existing evaluation approaches for styled text generation and propose an evaluation score specifically sensitive to calligraphic and typographic style, addressing a systematic flaw in the use of generic image generation metrics for this task. We then introduce methods for controllable handwritten and typewritten styled text generation, conditioning on content, style, and geometric variation. For visual content with extreme aspect ratios, such as panoramas, banners, and wide illustrations, we adapt large-scale pretrained diffusion models in a training-free manner, coordinating multiple generation views to produce globally coherent outputs at arbitrary aspect ratios without introducing new trainable parameters. Additionally, we tackle illustrations, characterized by irregular shapes and semi-transparent elements, and propose a pipeline to create, from a textual prompt, transparency-aware (RGBA) assets ready for compositing in complex layouts. Finally, we address multi-instance infographic editing by introducing an attention-partitioning mechanism for flow matching models that disentangles concurrent edits across multiple localized regions, preventing attribute leakage and enabling single-pass editing. We accompany this contribution with the first large-scale benchmark for multi-instance infographic editing, comprising images with up to 285 simultaneous edit instances, an order of magnitude beyond prior benchmarks. With this thesis, we aim to help close the loop between analysis, generation, and editing, enabling new applications in archive digitization, dataset augmentation, and content creation for Visually Rich Documents, while showing that models, metrics, and benchmarks developed for one stage of the pipeline can directly support the others.
Multimodal Models for Understanding and Generating Visually-Rich Documents
20-lug-2026
CUCCHIARA, Rita
CASCIANELLI, Silvia
File in questo prodotto:
File Dimensione Formato  
Quattrini.pdf

Open access

Descrizione: Quattrini.Fabio.pdf
Tipologia: Tesi di dottorato
Dimensione 95.51 MB
Formato Adobe PDF
95.51 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1414330
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact