Scribe re-identification in historical manuscripts is a challenging retrieval problem due to severe visual variability, document degradation, and the scarcity of annotated data. Most existing approaches rely on convolutional neural networks and often struggle to generalize in few-shot settings that are common in real-world archival scenarios. In this work, we investigate the effectiveness of Siamese Vision Transformer architectures for scribe re-identification under extreme data constraints. We formulate the task as a metric learning problem and conduct a systematic empirical study comparing convolutional baselines with multiple Transformer-based backbones, including Vision Transformers and hierarchical variants, in a 10-shot learning regime. Our results show that lightweight Vision Transformer models consistently outperform convolutional networks and larger Transformer variants across standard retrieval metrics, highlighting the importance of model capacity and inductive bias in low-data document analysis tasks. We further analyze the impact of different training objectives and architectural choices through extensive ablation studies, providing insights into the behavior of Transformerbased models for historical document retrieval. These findings suggest that carefully designed Transformer architectures offer a promising direction for scribe re-identification and related document analysis problems in digital humanities applications.

An Empirical Study of Siamese Vision Transformers for Scribe Re-Identification

Colombi E.;Foresti G. L.
2026-01-01

Abstract

Scribe re-identification in historical manuscripts is a challenging retrieval problem due to severe visual variability, document degradation, and the scarcity of annotated data. Most existing approaches rely on convolutional neural networks and often struggle to generalize in few-shot settings that are common in real-world archival scenarios. In this work, we investigate the effectiveness of Siamese Vision Transformer architectures for scribe re-identification under extreme data constraints. We formulate the task as a metric learning problem and conduct a systematic empirical study comparing convolutional baselines with multiple Transformer-based backbones, including Vision Transformers and hierarchical variants, in a 10-shot learning regime. Our results show that lightweight Vision Transformer models consistently outperform convolutional networks and larger Transformer variants across standard retrieval metrics, highlighting the importance of model capacity and inductive bias in low-data document analysis tasks. We further analyze the impact of different training objectives and architectural choices through extensive ablation studies, providing insights into the behavior of Transformerbased models for historical document retrieval. These findings suggest that carefully designed Transformer architectures offer a promising direction for scribe re-identification and related document analysis problems in digital humanities applications.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11390/1338430
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact