With the rising amount of real-time data in autonomous driving, surveillance systems, and medical imaging, streaming data analysis is becoming increasingly important but is severely constrained by concept drift and label scarcity. Current drift detection models are not easily adaptable, robust, or efficient in semi-supervised and high dimensional environments. To address these drawbacks, we suggest a novel adaptive concept drift detection framework for data streams based on Vision Transformers (ViT), called ViT-Drift. It uses a pre-trained ViT to extract high-dimensional patch-based embeddings, enabling sensitive monitoring of distributional changes across sliding windows. However, in contrast to traditional supervised learning techniques, the key to ViT-Drift lies in its pseudo-labeling mechanism, which uses entropy to dynamically label unlabeled data, minimizing reliance on manual labels. To guarantee model stability under distributional shift, an adaptive transfer learning strategy is introduced, which is based on lightweight fine-tuning of the model by incorporating a mixture of labeled and pseudo-labeled samples. In addition, consistency regularization is built in to make the model more robust in the presence of uncertainty and reduce the effect of noisy pseudo labels. Drift monitoring is performed through a responsive outlier detection mechanism combined with adaptive threshold adjustment. The experiments conducted on eight benchmark datasets demonstrate that on average, ViT-Drift can achieve an accuracy of 96.07%, surpassing Adaptive Windowing (ADWIN) (91.12%) and Drift Detection Method (DDM) (89.13%). It also shows lower errors on advanced metrics like Mean Absolute Error (MAE) (as low as 0.30), Root Mean Square Error (RMSE) (0.39) and Log Loss (0.41). It is stable, with less than 1.5% accuracy difference between folds, in the process of k-fold cross validation, and validated by the ablation study on the necessity of each module. The proposed approach is encouraging and can be adapted to detect concept drift for the high-dimensional visual streaming problems. Though the experiment shows the effectiveness of the proposed strategy for several benchmark datasets, further investigations are required to evaluate the effectiveness of the deployment, resource usage, and actual performance in an operational edge system.

A semi-supervised Vision Transformer framework for adaptive drift detection in data streams using transfer learning

Cascone, Lucia
2026

Abstract

With the rising amount of real-time data in autonomous driving, surveillance systems, and medical imaging, streaming data analysis is becoming increasingly important but is severely constrained by concept drift and label scarcity. Current drift detection models are not easily adaptable, robust, or efficient in semi-supervised and high dimensional environments. To address these drawbacks, we suggest a novel adaptive concept drift detection framework for data streams based on Vision Transformers (ViT), called ViT-Drift. It uses a pre-trained ViT to extract high-dimensional patch-based embeddings, enabling sensitive monitoring of distributional changes across sliding windows. However, in contrast to traditional supervised learning techniques, the key to ViT-Drift lies in its pseudo-labeling mechanism, which uses entropy to dynamically label unlabeled data, minimizing reliance on manual labels. To guarantee model stability under distributional shift, an adaptive transfer learning strategy is introduced, which is based on lightweight fine-tuning of the model by incorporating a mixture of labeled and pseudo-labeled samples. In addition, consistency regularization is built in to make the model more robust in the presence of uncertainty and reduce the effect of noisy pseudo labels. Drift monitoring is performed through a responsive outlier detection mechanism combined with adaptive threshold adjustment. The experiments conducted on eight benchmark datasets demonstrate that on average, ViT-Drift can achieve an accuracy of 96.07%, surpassing Adaptive Windowing (ADWIN) (91.12%) and Drift Detection Method (DDM) (89.13%). It also shows lower errors on advanced metrics like Mean Absolute Error (MAE) (as low as 0.30), Root Mean Square Error (RMSE) (0.39) and Log Loss (0.41). It is stable, with less than 1.5% accuracy difference between folds, in the process of k-fold cross validation, and validated by the ablation study on the necessity of each module. The proposed approach is encouraging and can be adapted to detect concept drift for the high-dimensional visual streaming problems. Though the experiment shows the effectiveness of the proposed strategy for several benchmark datasets, further investigations are required to evaluate the effectiveness of the deployment, resource usage, and actual performance in an operational edge system.
2026
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11386/4956655
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact