Test input generators (TIGs) are widely used to assess the robustness of Deep Learning (DL) image classifiers, yet they often produce invalid inputs that fall outside the semantic domain of the task, misleading quality assessment. While several automated validators have been proposed, there is a critical mismatch between automated and human validation criteria and, thus, automated validators are merely a proxy of domain validity, as perceived by human testers. We introduce DeepNaqqal, a supervised test input validator that learns validity directly from human-annotated labels using transfer learning on deep vision models. Our empirical study on automated validation of misclassification-inducing inputs compares DeepNaqqal against six state-of-the-art validators across three image classification tasks and multiple TIG families, using independent human assessment as ground truth. Our results show that DeepNaqqal consistently achieves the highest agreement with human judgments, while generalizing to unseen TIGs and remaining effective with substantially reduced labeled data.

DeepNaqqal: Human-Aligned Automated Validation of Test Inputs for Deep Learning

Riccio V.
2026-01-01

Abstract

Test input generators (TIGs) are widely used to assess the robustness of Deep Learning (DL) image classifiers, yet they often produce invalid inputs that fall outside the semantic domain of the task, misleading quality assessment. While several automated validators have been proposed, there is a critical mismatch between automated and human validation criteria and, thus, automated validators are merely a proxy of domain validity, as perceived by human testers. We introduce DeepNaqqal, a supervised test input validator that learns validity directly from human-annotated labels using transfer learning on deep vision models. Our empirical study on automated validation of misclassification-inducing inputs compares DeepNaqqal against six state-of-the-art validators across three image classification tasks and multiple TIG families, using independent human assessment as ground truth. Our results show that DeepNaqqal consistently achieves the highest agreement with human judgments, while generalizing to unseen TIGs and remaining effective with substantially reduced labeled data.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11390/1338428
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 1
  • ???jsp.display-item.citation.isi??? ND
social impact