Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack DetectionExplained for Beginners
Guray Ozgur, Fadi Boutros, Naser Damer
Abstract
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.
The Problem
For years, the field of face presentation attack detection (PAD)—commonly called "face anti-spoofing"—has treated itself as a uniquely facial problem. If you wanted to build a system that spots a fake fingerprint, a printed photo, or a video replay playing on a screen, the standard assumption was that you needed a dataset full of human faces. The logic followed that domain shift—changes in subject identity, demographics, lighting, and backgrounds—was the primary reason models failed when moved from one dataset to another. This framing has practical, heavy consequences: it drives the collection of massive biometric databases, raising consent and privacy concerns, and it forces researchers to design models that explicitly look for facial features.
In this paper, Fadi Boutros, Guray Ozgur, and Naser Damer challenge this fundamental assumption. They pose a provocative question: Does a PAD model actually need to see faces to do its job? Their core argument is that the visual artifacts introduced by the attack—the "spoof"—are not inherently tied to faces. Whether you are printing a photo of a face or replaying a video of one, the physical process of printing, displaying, and recapturing with a camera introduces specific optical and digital artifacts (like moiré patterns, pixelation, and gamma distortion). These "presentation attack instruments" leave traces on any object they touch. If this is true, then a model could potentially learn to detect these traces using any object, doing away with the need for face data entirely, and by extension, doing away with the privacy risks of collecting face data.
How It Works (The Technical Mechanics)
The authors introduce TPO, a cleverly constructed dataset designed to isolate the "presentation process" from the "facial content." TPO stands for Tomatoes, Potatoes, and Onions. The dataset consists of recordings of these vegetables—almost randomly chosen, 26 of each—under conditions that mimic standard face PAD protocols.
Here is the setup: The researchers captured bona fide (real) videos and images of the vegetables. They then created print attacks by printing the vegetable images on paper and recapturing them via camera. They created replay attacks by displaying the vegetable videos on a tablet or phone and recording the screen with another camera. Crucially, they used two different devices (a Microsoft Surface tablet and a Samsung Galaxy smartphone) and varied the distance (close and far), creating a "factorial protocol" that mirrors the variability found in real face datasets.
The technical heart of the experiment uses a FoundPAD architecture. This is a foundation-model-based approach: specifically, a CLIP ViT-B/16 image encoder with rank-stabilized LoRA (Low-Rank Adaptation) tuning on the query and value projections, topped with a linear two-class head. The CLIP backbone is frozen; only the LoRA weights and the head are trained. This is a parameter-efficient way to adapt a general-purpose vision model to a specific task.
The authors make a point about preprocessing: they use CLIP's native channel mean and standard deviation for normalization, rather than the typical ImageNet normalization. This small but vital detail ensures the model operates under the same statistical conditions it was pre-trained on.
The "mechanic" of the paper’s argument is this: By training a detector exclusively on tomatoes, potatoes, and onions—never showing it a human face—they test whether the model learns the spoof patterns or the face patterns. Because the vegetables have different shapes, colors, and textures, any successful transfer to face PAD would imply the model has learned something about the process of being printed or replayed, rather than the specific semantics of a nose or an eye.
Key Results & Benchmarks
The results are striking and form the technical backbone of the paper's claim.
- Face-Free Training is Surprisingly Strong: A detector trained exclusively on TPO (vegetables only) achieves an average AUC (Area Under the Curve) of 92.70% across four standard face PAD benchmarks (MSU-MFSD, CASIA-FASD, Idiap Replay-Attack, and OULU-NPU). This is the "MCIO" average.
- Outperforming Synthetic Faces: This TPO-trained model beats the same architecture trained on SynthASpoof, a dataset of synthetic faces specifically designed for PAD training. SynthASpoof only achieves 81.02% average AUC. This suggests that realistic presentation artifacts (from printing/replaying vegetables) are more informative for the model than rendered synthetic faces.
- Competitive with Real Face Data: The TPO-trained model remains competitive with models trained on real face datasets. In some protocols, it surpasses a single real face dataset; in others, it comes within one AUC point of the performance achieved by real face training.
- The Reverse Transfer: Interestingly, models trained on real face datasets transfer to TPO with AUCs ranging from 58.89% to 89.88%. While lower than the face-free transfer, these scores are consistently above chance (50%), suggesting that face models do learn some general "presentation process" cues, but they are heavily entangled with facial identity.
Perhaps the most compelling result regarding the "complementary" nature of the data is found in Table 3 of the paper. When the authors replaced part of a fixed face-training budget with TPO data (keeping the total computational budget the same), performance improved. The average AUC jumped from 89.33% (face only) to 92.55% (face + TPO), while the HTER (Half Total Error Rate) dropped from 17.54% to 13.97%. This indicates that the vegetable data provides "complementary information"—it teaches the model something about spoofing that the face data wasn't fully covering—rather than just being "more of the same" data.
Why It Matters (Key Takeaways)
This paper shifts the Overton window for how we think about biometric security and privacy. Here are the four key takeaways:
- The Face is Incidental: The most radical finding is that a PAD model can be trained on vegetables and generalize to human faces with near state-of-the-art performance. This suggests that what the model is "looking for" is not a face, but the physical traces of the recapture process. For the industry, this is a proof of concept that we can build effective anti-spoofing without relying on face images.
- Privacy by Design: If PAD models don't need faces, the pressure to collect massive, sensitive facial databases diminishes. This opens the door to "privacy-preserving" PAD systems that could potentially be trained on everyday objects or synthetic data, aligning biometric security with data minimization principles (a core concept in GDPR and similar regulations).
- Presentation Artifacts over Semantics: The finding that TPO outperforms SynthASpoof (synthetic faces) highlights that the instrument of the attack—the print texture, the screen refresh rate, the lighting—is the key signal, not the identity of the subject. This encourages researchers to focus on robustness to the recapture pipeline rather than chasing ever-more complex facial feature extractors.
- The "Tomato Principle": The paper establishes that diversity of presentation attack instruments matters more than the volume of data from a single source. Training on just one type of attack (print or replay) drops performance significantly, but combining them yields the best results. Furthermore, the "redundancy" of video frames is low value; a few diverse views are worth more than many repeated frames of the same vegetable. This has practical implications for how we collect and annotate future PAD datasets.
Ultimately, "Tomatoes, Potatoes, and Onions" is a landmark paper because it provides empirical evidence that face PAD is, at its core, a problem of detecting recapture artifacts. It proves that we can divorce the task of spoof detection from the ethics and logistics of facial data collection, provided we have a foundation model capable of learning the shared visual language of printing and replaying.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →