Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors

Sahabandu, Dinuka; Xu, Xiaojun; Rajabi, Arezoo; Niu, Luyao; Ramasubramanian, Bhaskar; Li, Bo; Poovendran, Radha

Computer Science > Cryptography and Security

arXiv:2402.08695 (cs)

[Submitted on 12 Feb 2024]

Title:Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors

Authors:Dinuka Sahabandu, Xiaojun Xu, Arezoo Rajabi, Luyao Niu, Bhaskar Ramasubramanian, Bo Li, Radha Poovendran

View PDF HTML (experimental)

Abstract:We propose and analyze an adaptive adversary that can retrain a Trojaned DNN and is also aware of SOTA output-based Trojaned model detectors. We show that such an adversary can ensure (1) high accuracy on both trigger-embedded and clean samples and (2) bypass detection. Our approach is based on an observation that the high dimensionality of the DNN parameters provides sufficient degrees of freedom to simultaneously achieve these objectives. We also enable SOTA detectors to be adaptive by allowing retraining to recalibrate their parameters, thus modeling a co-evolution of parameters of a Trojaned model and detectors. We then show that this co-evolution can be modeled as an iterative game, and prove that the resulting (optimal) solution of this interactive game leads to the adversary successfully achieving the above objectives. In addition, we provide a greedy algorithm for the adversary to select a minimum number of input samples for embedding triggers. We show that for cross-entropy or log-likelihood loss functions used by the DNNs, the greedy algorithm provides provable guarantees on the needed number of trigger-embedded input samples. Extensive experiments on four diverse datasets -- MNIST, CIFAR-10, CIFAR-100, and SpeechCommand -- reveal that the adversary effectively evades four SOTA output-based Trojaned model detectors: MNTD, NeuralCleanse, STRIP, and TABOR.

Subjects:	Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as:	arXiv:2402.08695 [cs.CR]
	(or arXiv:2402.08695v1 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2402.08695

Submission history

From: Bhaskar Ramasubramanian [view email]
[v1] Mon, 12 Feb 2024 20:14:46 UTC (6,900 KB)

Computer Science > Cryptography and Security

Title:Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators