Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors predominantly support global or single-instruction edits and struggle with multi-instance scenarios, where multiple parts of a reference input must be edited independently without semantic interference. We identify this limitation as a consequence of globally conditioned velocity fields and joint attention mechanisms, which entangle concurrent edits. To address this issue, we introduce Instance-Disentangled Attention, a mechanism that partitions joint attention operations, enforcing binding between instance-specific textual instructions and spatial regions during velocity field estimation. We evaluate our approach on both natural image editing and a newly introduced benchmark of text-dense infographics with region-level editing instructions. Experimental results demonstrate that our approach promotes edit disentanglement and locality while preserving global output coherence, enabling single-pass, instance-level editing.

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing / Zaccagnino, C., Quattrini, F., Simsar, E., Tintoré Gazulla, M., Cucchiara, R., Tonioni, A., Cascianelli, S.. - (2026). (43nd International Conference on Machine Learning, ICML 2026 Seoul, Korea July 6th - 11th, 2026).

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Carmine Zaccagnino
;
Fabio Quattrini;Rita Cucchiara;Silvia Cascianelli
2026

Abstract

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors predominantly support global or single-instruction edits and struggle with multi-instance scenarios, where multiple parts of a reference input must be edited independently without semantic interference. We identify this limitation as a consequence of globally conditioned velocity fields and joint attention mechanisms, which entangle concurrent edits. To address this issue, we introduce Instance-Disentangled Attention, a mechanism that partitions joint attention operations, enforcing binding between instance-specific textual instructions and spatial regions during velocity field estimation. We evaluate our approach on both natural image editing and a newly introduced benchmark of text-dense infographics with region-level editing instructions. Experimental results demonstrate that our approach promotes edit disentanglement and locality while preserving global output coherence, enabling single-pass, instance-level editing.
2026
9-feb-2026
43nd International Conference on Machine Learning, ICML 2026
Seoul, Korea
July 6th - 11th, 2026
Zaccagnino, Carmine; Quattrini, Fabio; Simsar, Enis; Tintoré Gazulla, Marta; Cucchiara, Rita; Tonioni, Alessio; Cascianelli, Silvia
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing / Zaccagnino, C., Quattrini, F., Simsar, E., Tintoré Gazulla, M., Cucchiara, R., Tonioni, A., Cascianelli, S.. - (2026). (43nd International Conference on Machine Learning, ICML 2026 Seoul, Korea July 6th - 11th, 2026).
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1414430
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact