Ali Samir

← The writing map

Deep Diaries · · 2 min read

Evil Model

Bringing Neural Network Models into Malware Attacks

Originally published on Deep Diaries on Substack. Reproduced here as written.

Figure 1 from “Evil Model”

Recently, I've found myself deeply intrigued by a subject that holds significant importance and captivation, particularly for individuals like me with a background in cybersecurity. The focal point of this topic revolves around the covert use of neural networks for delivering malware to a target device, all while evading detection by antivirus software. Within this realm, researchers have delved into the intricacies of concealing malware within the parameters of machine learning models, drawing inspiration from the art of steganography.

  1. Least Significant Bit (LSB): In the LSB approach, the attacker meticulously dissects their malware into individual bits and strategically inserts these bits into the least significant bit of the network's parameters. However, this technique necessitates extensive alterations to the parameters to accommodate the entire malware, effectively tethering the endeavor to the network's dimensions.

  2. Resilience Training: An alternate strategy relies on the model's innate ability to recover from errors through training. Here, the attacker orchestrates a complete overhaul of parameter values to mirror those of their malware, locking these parameters in place during training to prevent any inadvertent alterations. Subsequently, standard training procedures are carried out.

  3. Value Mapping: The value mapping approach entails the meticulous identification of parameter values closely mirroring those of the concealed malware. This mapping endeavor minimizes the requirement for sweeping modifications but demands access to an extensive mapping database.

  4. Sign Mapping: Parallel to value mapping, sign mapping concentrates solely on the mapping of sign bits rather than the entire value.

  5. Float32 Manipulation: Some parameters are represented as float32 numbers conforming to the IEEE standard. By manipulating the mantissa (M) while keeping the exponent (E) fixed, subtle variations in the actual value can be introduced. The attacker then manipulates the mantissa (M) to embed their malware within the parameter.

Live Demo:

Using Metasploit payload injected into a simple PyTorch model trained on MNIST dataset

Naturally, the attacker must establish explicit rules for embedding, extraction, and triggers based on specific conditions. This intricate orchestration empowers them to distribute the malware as an update to all users of the model or even release it on the internet as a pre-trained model, all while avoiding significant increases in size or detrimental effects on accuracy. This model could potentially breach security gateways, operating incognito.

Paper:

[1] https://arxiv.org/abs/2107.08590

[2] https://arxiv.org/abs/2109.04344

First Publish: Post on LinkedIn