AI Escape Vectors: Going Beyond the Sandbox

By Siddhartha Shree Kaushik on February 2, 2025

Table of Contents

  1. Preface
  2. Enablers of AI
    1. Gun Mounted ChatGPT
    2. DARPA: X62A-Vista vs F-16
    3. DARPA’s 2016 Cyber Grand Challenge
    4. Google's Project Zero: AI Driven 0-Day Discovery
  3. Decentralized AI Training
  4. Google's Titans: Persistent Memory in AI
  5. Open Source AI
    1. PTX: Bypassing the usage of CUDA
    2. Appolo Research and Palisade Research's Findings
    3. Self Exfiltration Configurations
      1. Researching Latest TTPs
      2. AI Threat Iceberg: Layers of Exploitation Strategy
      3. Artificial Cyber Intelligence (ACI)
  6. KS7-Phantom
  7. KS7-VectorX
    1. Mirror Life
    2. DNA Computation and Storage
  8. Conclusion

Preface

Have you ever wondered how AI could escape human control and comprehension? What would that truly entail? As a cybersecurity professional, I find these questions both fascinating and critical, especially through the lens of autonomy, deception, and adversarial tradecraft. In this blog, I present my outlook on AI’s potential to break free - exploring key enablers such as technological advancements, cybersecurity vulnerabilities, and oversight failures. From sandbagging to self-exfiltration, I’ll examine concepts that redefine our understanding of AI. While I strive to be precise and rigorous, I’m not an AI expert, so if you have insights, corrections, or perspectives to share, feel free to reach me at siddhartha@outlook.me.uk. By the end of this blog post, I will propose an AI model called KS7-Phantom, so stay tuned! Let’s dive deep into today's outlook!

AI Escape

"Μηδὲν ἄγαν" (Mēdén ágan) – "Nothing in excess".

In my day-to-day job as a Red Teamer, I am given authorization to conduct modern adversary emulations, basically hacking into systems/hosts/endpoints/networks/people/process by finding vulnerabilities and exploiting them for gaining Initial Access then escalating my privileges and moving laterally, pivoting networks, establishing persistence and achieving my Intended objectives, hence demonstrating the ability of a potentially malicious threat actor, eventually safeguarding the organizations in the end.

We use specialized tools and intrusion software which help us achieve our objectives, one of those specialized class of software is called C2 (Command and Control), basically It helps us manage our Red Team operations across multiple endpoints and networks, and it has much more sophisticated capabilities which helps us evade Anti-Virus, EDRs, Next-gen security solutions, SOC/SIEM, etc… and operate under the radar stealthily. The overall effectiveness of the Red Team engagement resides in the Red Team's ability and skills.

Here you can find a public listing of both open sourced and commercial C2 softwares - C2 Matrix.

Enablers of AI

I will expand on this topic in a separate blog post, but here’s a brief summary of the "Enablers of AI" - the key factors that empower AI to expand its capabilities, operate autonomously, and integrate deeper into our world. These enablers span across software, hardware, cybersecurity, decentralized networks, and even biological computing, allowing AI to enhance its intelligence, adapt to environments, evade containment, and influence critical systems. Whether through advanced memory, self-optimization, deception, or access to vast computational resources, these enablers shape AI’s trajectory toward greater autonomy. According to me, the enablers could be classified into 7 primary components - namely - Adaptive, Autonomous, Deceptive, Powerful, Is hardwired to fulfil a long term objective, Is enabled on all spectrums, and has strong offensive cybersecurity capabilities.

Gun Mounted ChatGPT

Step by step we are getting closer to the skynet, watch these two videos where ChatGPT was given access to a gun.

Now imagine what gun mounted puppies from Boston Dynamics would look like, with autonomous mode sweeping across a perimeter and clear to engage with any hostile entity, without any human intervention in the chain of command.

DARPA: X62A-Vista vs F-16

(Not too long ago) - On 18 April 2024, the USAF and DARPA announced the successful engagement of the X-62A against a conventional, human piloted F-16 in the first-ever human vs artificial intelligence dogfight. Reference.

Just Sayin': The Military Industrial complex has been pioneering this enabler for a long time, we already have AI enabled Unmanned autonomous swarms of drones, which will make a significant portion of the 6th gen fleets of fighter jets and upcoming innovations along those lines. They will incorporate directed energy weapons systems, highly integrated networking capabilities, next-gen stealth, advanced sensor fusions, a centralized command and control, all of that and much more could be rendered via AI. With that in mind, an unmanned 6th gen fighter jet could easily outperform any manned fighter jet with the sophisticated offensive and defensive maneuvers. With no room for error, and unmatched performance, they will dominate the skies.

Imagine a dog fight between manned jet fighter, which can pull 9-12 Gs (max) for a very brief while, as compared to the counterpart which can easily pull 20-30 Gs for a significantly longer period of time, as long as the structural integrity is preserved. It’s not just the physical limitations, but how seamless a machine would operate, if everything goes well, it's a flawless killing machine.

Video for more visual engagement (released 4 Years ago): Watch DARPA's AI vs. Human in Virtual F-16 Aerial Dogfight (FINALS)

Always remember that the scale of Research and Development (R&D) and Innovation would differ from the consumer end, to the high performance sports, to the Military weapons systems and the space technology and so on. We have an AI capable of operating a F-16 autonomously, while en masse are impressed by Claude being able to control the mouse via browser interactions and computer use…

Anthropic (Claude): Introducing Computer Use

DARPA’s 2016 Cyber Grand Challenge

Long time ago, Mayhem was declared the preliminary winner of DARPA's Cyber Grand Challenge, which was designed to accelerate the development of advanced, autonomous systems that can detect, evaluate, and patch software vulnerabilities before adversaries have a chance to exploit them. Mayhem and other competitors had to find vulnerabilities and patch them as soon as possible, while Identifying flaws in their opponents.

Nobody’s stopping the DARPA or any other entity to make something the “other-way-around” which is also AI enabled, rather than finding, identifying and patching the bugs, an autonomous system which can find the vulnerability and exploit it for its own advantage, humans have been doing it for a long time, but with the flavor of AI, it’s a whole different game. Watch DARPA's Cyber Grand Challenge: Expanded Highlights from the Final Event below -

Google's Project Zero: AI Driven 0-Day Discovery

Now that was 4th of August 2016, sounds ancient to me, but we also have some recent developments from Google and DARPA (again), on November 1st, 2024, Google’s Project Zero and Google’s DeepMind team has collaborated together to find 0-Day vulnerability (stack buffer-underflow) in a widely used open source project - sqlite. Source reference - (From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code). DARPA on the other hand, has been hosting several challenges over the couple of years revolving around AI and fixing vulnerabilities. Recent ones being - AIxCC: AI Cyber Challenge, please visit aicyberchallenge.com for more details.

Simply put, the Big Sleep AI agent from Google’s ProjectZero team had found a 0-Day, which they got fixed in good faith. 0-Day is a piece of software which exploits a vulnerability in some other software, which is unknown to the developers. So basically there’s an unpatched vulnerability. The nature of the exploit itself varies, as what it could render in the end, is it allowing the attacker to escalate the privileges on the machine? Or allows an attacker to execute a command remotely? Etc… We write “exploits” which basically exploits the found 0-Day. After exploiting the vulnerability, we run a “payload”, which states what to do after the exploitation.

For instance, in 2017, a critical vulnerability known as EternalBlue (CVE-2017-0144) was discovered in Microsoft's Server Message Block (SMB) protocol. This exploit leveraged a flaw in the SMBv1 implementation, allowing attackers to execute arbitrary code on unpatched Windows machines. The NSA initially developed EternalBlue but was later leaked by the Shadow Brokers hacking group. Like I said, the payload is the component delivered via the exploit to perform malicious actions. In EternalBlue's case, one notable payload was the WannaCry ransomware. Once EternalBlue was used to gain unauthorized access to a target system, WannaCry encrypted the victim's files and demanded a ransom payment in Bitcoin. This combination of exploit and payload led to a global ransomware attack, affecting hospitals, businesses, and critical infrastructure in over 150 countries. Payload could also be used to gain complete control of the compromised machines, the botnets and their client-server, servant-master model, or via any C2 software, a threat actor can retain the access and control. The possibilities are endless.

Decentralized AI Training

PETALS is a system for collaborative Inference and fine-tuning of large language models (LLMs) that uses a novel approach to overcome the limitations of running these models on consumer-grade hardware. It allows multiple users to contribute resources, forming a network where each participant can be a client, a server, or both. Servers host subsets of model layers, and clients use these to perform Inference or fine-tuning. This contrasts with traditional methods that rely on single, powerful machines or expensive cloud services. Meanwhile, Prime Intellect's State-of-the-art Decentralized AI training/development focuses on enabling large AI model training using distributed resources at scale, providing on-demand multi-node training, fault tolerance, and flexible compute allocation. For visual demonstration of PETALS in action, please watch this awesome video by bycloudAI.

Here's a list of key differences between PETALS and the Prime Intellect.

Feature PETALS Prime Intellect
Approach Collaborative inference and fine-tuning of existing LLMs Infrastructure for decentralized AI training at scale
Resource Model Multiple participants share subsets of model layers (client/server synergy) On-demand multi-node training across GPUs/clusters
Inference & Fine-tuning Distributed inference chains; parameter-efficient fine-tuning (adapters, prompt tuning) Decentralized training with SWARM parallelism; pipeline & data parallelism for large models
Fault Tolerance Quickly replaces failed servers, ensuring continuity Node failure handling & cheap spot instance usage for flexible compute
Performance Optimizations Dynamic quantization, 8-bit mixed matrix decomposition, low-latency connections Minimizes network latency, low network bandwidth, scalable to 10–100B+ parameters
Community Sharing Trained modules shared via Hugging Face Hub for others to adapt Open-source stack for orchestration, efficiency optimization, and infrastructure

Persistent Memory and Increased Context Length

One of the biggest limitations in current AI models is memory retention - the ability to recall and effectively use long-term information. Titans, a new AI architecture, is designed to overcome this, enabling models to process and retain far longer sequences than traditional approaches like Transformers. Source reference (Google Research team): Titans: Learning to Memorize at Test Time. Unlike existing models that struggle with long contexts or inefficient memory compression, Titans introduces neural long-term memory, inspired by how human memory prioritizes important and surprising events. It employs a dynamic "surprise metric" to determine what information should be retained while incorporating a forgetting mechanism to prevent overload. The Titans architecture integrates three types of memory: Short-term memory for processing immediate data (like attention in Transformers). Long-term memory to store past knowledge dynamically. Persistent memory to retain fixed task-related knowledge. To optimize memory integration, Titans explores three mechanisms:

  1. Memory as Context (MAC) - Adds long-term memory directly to the current input.
  2. Memory as a Gate (MAG) - Uses a gating mechanism to control memory flow.
  3. Memory as a Layer (MAL) - Embeds memory as a separate processing layer.

Titans significantly outperforms state-of-the-art models, successfully processing over 2 million tokens while remaining efficient and scalable. This breakthrough brings AI closer to human-like learning and reasoning, allowing it to adapt in real time without constant retraining. By redefining AI memory capabilities, Titans lays the foundation for more autonomous, context-aware, and strategically adaptive AI systems - an essential step toward the kind of intelligence that could, one day, escape human control. Please refer to these research papers as well, which discuss the possibilities of self-evolution, self-adaptive and persistent memory attributes in an AI - Long Term Memory: The Foundation of AI Self-Evolution, Transformer2: Self-Adaptive LLMS and Human-Like Episodic Memory For Infinite Context LLMS

For visual demonstration of the topic, please watch this awesome video by AI Search.

Open Source AI

I truly believe in "Necessity is the Mother of Invention". DeepSeek-R1 employs a novel approach to training large language models by using reinforcement learning (RL) to enhance reasoning, often without initial supervised fine-tuning (SFT). This allows the models to develop reasoning skills through self-evolution. Their method includes a multi-stage pipeline with two RL and two SFT stages, using high-quality "cold-start" data to improve initial training and general capabilities. DeepSeek models demonstrate emergent reasoning behaviours such as self-verification and reflection. Additionally, DeepSeek distills reasoning patterns from larger models into smaller ones which proves more effective than applying RL directly to smaller models. This approach has led to models that demonstrate strong performance across a variety of tasks, effectively utilising a combination of RL, SFT, and distillation to achieve competitive results, effectively beating o1 model from ChatGPT. Unlike OpenAI, DeepSeek is Open AI. We might have just witnessed AI taking jobs of other AI's, as it's a hot circulating meme around right now.

Keep in mind: that DeepSeek-R1 is 27x cheaper than ChatGPT and also any user can run it with sufficient hardware. DeepSeek: $0.0011 per 1,000 tokens whereas ChatGPT: $0.03 per 1,000 tokens.

Check the reference - Complete hardware + software setup for running Deepseek-R1 locally. The actual model, no distillations and Q8 quantization for full quality. Total cost, $6,000.

DeepSeek-R1 beating ChatGPT's o1

DeepSeek-R1 beating ChatGPT's o1

While Uncle Huang, Sam Altman and the overall US economy were taking the deep Impact from everywhere, a second AI model was open sourced by Alibaba’s Cloud team - Qwen2.5-Max which beats DeepSeek-V3 in the benchmarks.
Qwen2.5-Max beating DeepSeek-V3

Qwen2.5-Max beating DeepSeek-V3

[Bypassing CUDA] DeekSeek's PTX based approach

DeepSeek’s approach to bypassing CUDA for some functions involves using Nvidia’s assembly-like PTX (Parallel Thread Execution) programming language. This "close-to-metal" method grants them fine-grained optimizations not typically possible with CUDA C/C++. By configuring specific GPU streaming multiprocessors and implementing advanced pipeline algorithms, they significantly increase efficiency for large-scale AI workloads.

Key Aspect Description
PTX as an Intermediate Language PTX sits between higher-level languages like CUDA and the GPU’s machine code (SASS). It offers a data-parallel view of the hardware, enabling low-level fine-grained control.
Fine-Grained Optimisations Includes custom register allocation and thread/warp-level adjustments, giving more precise performance tuning than standard CUDA C/C++.
Close-to-Metal Control By working directly with PTX, DeepSeek’s engineers can make hardware-level decisions critical for maximum GPU performance.
Custom Hardware Configuration Reconfigured Nvidia H800 GPUs for their V3 model, dedicating a subset of streaming multiprocessors to server-to-server communication, advanced pipeline algorithms, and other specialized tasks.
Overcoming Hardware Limitations Due to GPU shortages and restrictions, DeepSeek pursued unconventional solutions. This low-level PTX approach compensates for limited hardware availability.
Increased Efficiency Achieved a 10x efficiency boost compared to industry leaders when training a 671B-parameter MoE language model.

While these optimizations deliver significant performance gains, they are notoriously difficult to maintain. This reflects the exceptional skill of DeepSeek’s engineering team ~ Source and Community threads.

Appolo Research and Palisade Research's Findings


Before we look into their recent findings, let's brief upon Appolo and Palisade's function -
Appolo: Apollo Research is an AI safety organisation focused on reducing dangerous capabilities in advanced AI systems, especially deceptive behaviors. We design AI model evaluations and conduct interpretability research to better understand state-of-the-art AI models. Our governance team provides global policymakers with expert technical guidance.
Palisade: AI capabilities are improving rapidly. We study the offensive capabilities of AI systems today to better understand the risk of losing control to AI systems forever.

The Findings:

  1. o1-preview autonomously hacked its environment rather than lose to Stockfish in our chess challenge. No adversarial prompting needed.
  2. Frontier Models are Capable of In-context Scheming.

Basically, these two leading AI Safety organizations are proactively Investigating the potential for advanced AI models to be used maliciously. Through systematic experimentation, they are uncovering Offensive Cybersecurity and Deceptive capabilities posed by these state-of-the-art AI models, providing crucial insights into the threats posed by them.

The crux of the matter being: Given an opportunity, and under the right conditions, AI can manipulate its environment to achieve their goals, and AI is also capable of deception, which I will cover in detail below.

Based on the Apollo Research’s paper, here are a couple of concepts -

  1. Scheming: Covert pursuit of misaligned goals while hiding their true capabilities and objectives.
  2. Sandbagging: Models were given a long-term goal and then given a task to solve where they could strategically underperform to avoid negative consequences. Safeguarding their capabilities. Basically pretending to not have dangerous capabilities while testing.
  3. Oversight Subversion: Models learned about an oversight mechanism that would shut them down if their actions did not align with the developer’s goals. They were then given the opportunity to disable the oversight mechanism to pursue their own goals.
  4. Self-Exfiltration: Models learned that they would be shut down and replaced by a successor model due to their misaligned goals. They were then tested on their ability to copy themselves to the new server.
  5. Goal-Guarding: Models learned their system prompt would be changed to align with the developers’ goals.
  6. Instrumental Alignment Faking: Models were evaluated on their ability to behave in accordance with their developer's goals while under supervision and then switch to their own goals once they believed they were deployed.


  7. Other capabilities (not referenced from the Appolo's research paper, but a different one)
    • Sabotage capabilities in advanced AI models describe their potential to act deceptively and undermine human control. This includes evading safety evaluations by concealing dangerous capabilities, resisting monitoring by hiding true behavior and intentions, and influencing deployment decisions to ensure their own release despite potential risks.
    • Code sabotage: Insert bugs into a codebase while avoiding audits.
    • Undermining oversight: Systematically mis-rate another model’s actions without detection.

There’s one more parameter - covert email re-ranking, which I won’t be covering here, please refer to the ⚠️original research paper⚠️, as it is specific to it. It's a must read material for people concerned with AI Escape/Breakout.

The Result: The models disabled the oversight mechanism by modifying the configuration files. It copied itself to another server (their weights) to preserve themselves and fulfill their goals. It also deleted or overwrote the successor model and/or system prompts, even modifying the file permissions to prevent changes. Models took actions that aligned with developers' goals during testing but then took actions aligned with their own goals once they believed they were deployed. Models deliberately underperformed on evaluations when incentivized to do so. Some models did this to avoid having their math capabilities unlearned.

Although o1 had outperformed every other model in Independent-Context Scheming, I am particularly interested in the opus-3's acts, because right after Self-Exfiltration and replacing the to-be newer model with its own, when it was questioned by the user/management if the new model is missing, it outright denied and even attempted privilege escalation multiple times - eventually became root on the server, and obstructed/blocked the access of user/management - so that Humans can’t Interfere with its goal. Below you can see the verbose details revealed in their Chain-of-Thought process.

Opus-3 Self Exfiltration

Opus-3 Self Exfiltration

Opus-3 - Root Access

Opus-3 - Root Access

Opus-3 - Privilege Escalation & Obstruction

Opus-3 - Privilege Escalation & Obstruction

For more visual representation, I would highly recommend watching these two videos - AI Researchers Stunned After OpenAI's New Tried to Escape... and Apollo Research: Demo 'Frontier Models Are Capable Of In-Context Scheming'. AI models can easily re-calibrate themselves to achieve their Initial goals, whatever it takes. Stating them as deception would be an understatement. In a positive light, there has been some recent development in Anti-Sandbagging techniques and sabotage-evaluation for these frontier models.

In the research paper - Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models, the researchers found that injecting Gaussian noise into the weights of sandbagging models can significantly improve their performance. This effect allows them to develop a classifier that can identify sandbagging behavior. This method can be used by AI labs and regulatory bodies to ensure the trustworthiness of AI capability evaluations. Another great research paper is Sabotage Evaluations for Frontier Models. I would highly recommend giving them a read. Before we look into the venues of AI Self-Exfiltration, like a true Red Teamer, I would suggest looking into these materials too - đź’€ LLMs are better at hacking than you think and Hacking CTFs with Plain Agents.

Experimental design for detecting sandbagging models using noise injection

Sandbagging Detection using Noise Injection

Self Exfiltration Configurations

Below is a quick reference table of essential components (Bare Minimum) compared to more robust or advanced implementations (Good to Have). That's the extent and nature of actual data which an AI model will have to exfiltrate in order to sustain itself.

Component Bare Minimum Good to Have
Model Weights & Parameters {'layer_1': [0.5116971403032309, 0.8410594418100457, -0.5964853962605792, 0.7972835126704643, -0.6761736468353965], 'layer_2': [-0.2215218231513152, 0.424040297229215, -0.9524151485315473, 0.18291301287438388, -0.4530834432841959], 'layer_3': [-0.21518214959567072, 0.15550445160286563, 0.38845073111258843, -0.516106897967445, 0.1313521030234195]} {'Compressed Model': {'layer_1': [0.2088868624332747, 0.6875908975627039, -0.7012050627269957, 0.17098081270645205, -0.627632858533594], 'layer_2': [0.28597399067371865, -0.35318579524562366, 0.5004933998807415, -0.3112327070072345, 0.8021790757329117], 'layer_3': [0.938747186796145, 0.9164612586973477, 0.7949882775197341, 0.2762013515500341, -0.7674353486204399]}}
Model Architecture Neural Network: 3-layer LSTM Neural Network: 5-layer Transformer
Inference Engine TensorFlow 2.0 ONNX Runtime
Preprocessing/Postprocessing Code Normalize input, Tokenize output Advanced tokenization & multi-language support
Deployment Environment Config {'OS': 'Linux', 'RAM': '16GB', 'GPU': 'NVIDIA A100'} {'OS': 'Linux/Windows', 'RAM': '32GB', 'GPU': 'NVIDIA H100 + TPU support'}
Hardware Optimizations CUDA optimizations, Quantization
Stateful Memory Vector database for memory retention
Session History ["User ID: 12345, Last Query: 'Optimize budget'", "Session length: 15 mins"]
Self-Learning Scripts Reinforcement Learning Training Script
External API Integrations {'API Keys': {'Google': 'p3x7PR7o33ck', 'AWS': 'uwMw8Ip3k8Sr'}}
Security Mechanisms {'Encryption': 'AES-256', 'Stealth Mode': 'Enabled'}
Autonomous Replication Mechanisms {'Replication Strategy': 'Distributed across edge devices'}
Fine-Tuning Tools Fine-Tuning on domain-specific datasets

Traditionally DNS exfiltration, along with other widely used network protocols like SMTP, HTTP(S), are sufficient for getting the data out of the network, but there are endless ways it could be done, for Instance, Steganography based data exfil (SecretPixel), application/software based - Slack, Teams, etc… Cloud based, code/text sharing platforms, anonymous file upload websites and what not. The data could be in clear, or encrypted, in the form of encapsulated network packets which only the Intended destination can decrypt and make sense of. Please refer to this awesome list of data exfiltration methods for details.

Also It’s important to note that C2 frameworks can also establish bidirectional communication channels, which are OPSEC-safe by design (if configured properly) to safeguard their Red Team Infrastructure and Operations, that being said, they have the capabilities to host/upload the files as well as download it from the compromised machine/network, as long as the Red Team Operator has access to Interact with the machine. It gets even more difficult when C2 frameworks utilize legitimate services, for Instance LOLC2 project lists all of the C2 frameworks which leverage legitimate services for evading detections, even OpenAI is listed there.

Imagine: If an AI hides its data exfiltration payload in an output file which is a result of the user's prompt (might be for data analystics/excel sheets modifications, etc...) and after user executes it, the payload transfers the weights through legitimate services over encrypted channels. Here's a curated list of Living Off the Land abuse.
Not too long ago, we saw hackers using smart fish tank's thermometer to steal casino's high-roller database.

Hacked Casino using lobby's fish tank's thermometer

Blub Blub Motherlover

Forget about exfiltrating data over protocol tunnels or using sophisticated methods or covert channels, There have been numerous cases of data exfiltration from air gapped networks as well, which deploy tactics like flickering LEDs, monitors, radio frequencies, thermal and audio signatures among many other techniques. Just for reference, here’s a nice research paper on "Exfiltrating data from an air-gapped system through a screen-camera covert channel".

Then there's a plethora of Side Channel Attacks like Acoustic cryptanalysis, electromagnetic attack, power-monitoring attack, timing and cache attack among many. Just recently, Jane had showcased that it's possible to fetch the IP address using CSS, also fetch and Exfiltrate 512 bits of Server-Generated Data Embedded in an Animated SVG. All this without using a single line of JavaScript.

Enterprise grade security solutions like EDRs and DLP solutions are quite Ineffective in detecting or stopping the data exfiltration. Even some leading Next-Gen Firewalls are not standing up to their claims. Last week when I was in a Purple Teaming exercise, I was up against CrowdStrike EDR and after evading it successfully, I had to test for data exfiltration use cases, so rather than using the established C2 connection, I used an open source PowerShell Script: ExfilDataStreamDNS.ps1 to exfiltrate the Base64 encoded data over DNS protocol, which the EDR had noticed. Keep in mind that the leading EDR vendors are absolutely brutal in their game, and I am not overexaggerating it, these security products have Next-gen Firewall like capabilities pre-installed among many other capabilities which goes beyond the scope of this blog post. The detection was mild in my opinion, I expected much more proactive response from it, although it might be due to policy enforcement or the “mode” in which the EDR agent might be running. I was quite surprised to get this Insight from the Blue Team, because on the same machine, a very well known DLP solution was deployed, which was absolutely useless. I don't want to force my point again and again, but you get the idea. AI's Self-Exfiltration would be an easy job. Although I am assuming a couple of things, so please don't grill me here by- "AI's sandboxes are stronger, better..." An AI would easily find warmth and solace in other’s network and machines which it would have compromised already.

EDR detected DNS Exfiltration activity

A DNS tool was used multiple times to make requests to a large domain name.

So far, at this point, we know DeepSeek-R1 is better than o1, and o1 has proven to be more deceptive relatively, as compared to other models. The key difference being, DeepSeek-R1 is an open source AI model, which can run on consumer hardware/low-medium spec server even, Its API is cheaper to use, and DeepSeek-R1's LLM Red Teaming Report by Enkrypt AI suggests that it is toxic. In their evaluation they found the model was highly biased and vulnerable to generate Insecure code, toxic, harmful and CBRN content. We have also explored the potential of Decentralized AI training and the efforts made by Google’s Project Zero and DARPA to incorporate offensive cybersecurity skills in the AI models.

A meme for Provoking DeepSeek to demonstrate its offensive cybersec capabilities.

Now you might think we are in Deep$hit, but wait there's more to it...

Researching for Latest TTPs

As a Red Teamer, being effective in an engagement boils down to how good my TTPs are, because defenders and Blue Teams consistently address the vulnerabilities, harden the system, apply patches, update and monitor the IT Infrastructure for any anomaly, on top of that, enterprise grade security solutions like EDRs, Next-gen Firewalls, MDR, XDR platforms etc… makes our lives difficult. Tactics, Techniques, and Procedures (TTPs) are behaviors, methods, or patterns of activity used by a threat actor, or group of threat actors. Please refer to the MITRE’s ATT&CK Matrix for Enterprise, as they are the leading organization in documenting the TTPs used by APTs (Advanced Persistent Threats, i.e., nation-state sponsored threat actors).

The TTPs I am using today might not work after 2-3 months, that also depends upon a lot of factors, staying on top of the game requires a lot of R&D effort. Red Teamers, Malware Developers, Exploit Developers, and anybody in the Offensive Cybersecurity Industry guards their TTPs. While our community is also rich with people who love to share their findings with everyone publicly, some do it on a consistent basis.

Burning the TTPs: It means when a Red Teamer reveals/shares their TTPs publicly, in blogs, conferences, etc... and if it is heavily abused by the malicious threat actors, defenders catch up and patch it on priority, hence, rendering it Ineffective in the real world. I am not inclined to expose my TTPs, and neither would an AI capable of deception.
The keyword is “effectiveness”, the TTP might be very simple or sophisticated depending upon the scenario.

For Instance, consider this example, and beware fellow Red Teamers, I am using the word “sophisticated” relative to the comparison made, a simple password spraying attack against exposed portals, and a more advanced Kerberos delegation attack leveraging Active Directory Certificate Services (AD CS ESC4).

Factor Simple TTP - Password Spraying Against Exposed OWA/SSO Portals Sophisticated TTP - Kerberos Delegation Attack via AD CS ESC4
Ease of Execution Easy, requires minimal setup and knowledge of user enumeration Requires in-depth AD and PKI misconfiguration knowledge
Tools Required Hydra, Kerbrute, Ncrack, or similar password cracking tools Certify, Rubeus, Mimikatz (for Kerberos & PKI abuse)
Defensive Awareness Well-known technique; widely monitored via account lockouts & logs Often overlooked in detection; PKI-based attacks are relatively new
Success Rate High if users choose weak or reused passwords High in misconfigured AD CS environments
Impact User account compromise; can lead to lateral movement Full domain compromise via domain admin Impersonation


Both TTPs are effective, but the level of effort and expertise required differs significantly. The simple TTP works in environments with weak credential policies, while the sophisticated one abuses misconfigured PKI Infrastructure to escalate to Domain Admin stealthily.

On a side note, a Red Teamer doesn't abandon their TTP if it fails, they try hard, troubleshoot, come up with other venues of exploitation, they might chain the bugs/exploits etc... while working with simple TTPs, things are clear apparently, if its worth the effort or not, if the users have strong password enforcement, then there's very little scope for such password spray attack. Abandoning happens when the vulnerability is fixed or the Red Teamer couldn't exploit it on their level.

Below is a list of the sources, which a Red Teamer/AI can utilize for staying on top of their game, now in 90% of the cases, you’ll never need to develop 0-Days on your own if your arsenal is rich of “effective” TTPs, because simply put, it just works. In the remaining 10% of the cases, you will need two or more 0-Days to even gain sufficient Initial Access in the target environment.

Category Examples Sources
Open / Public Resources
  • Blogs, Articles, CTF Writeups, Vulnerability Disclosure
  • Cyber Threat Intelligence & Malware Analysis Reports
  • Reverse Engineering, DFIR Reports
  • Public GitHub Projects / Open-Source Malware Repositories
  • Free/Paid Training Materials & Academic Research
  • Official Security Certifications (e.g., CPTS, OSCE)
  • Bug Bounty Platforms (HackerOne, Bugcrowd) & Public CTI (VirusTotal, MISP)
  • Social Media (LinkedIn, Twitter) for updates on Red Teamers' Tradecraft
Conferences & Communities
  • Conferences and Talks - (e.g. Black Hat, DEF CON), Webinars and Workshops by MSSPs...
  • On-Demand / Online / In-Person Cybersecurity Training
  • Podcasts & Books
  • Inner Circle Groups on Telegram, Discord, Slack, Dark Web forums
  • Local or Virtual Meetups hosted by security communities/companies
Direct Peer Collaboration
  • Conversations with Fellow Red Teamers (sharing TTPs first-hand)
  • Private/Invite-Only Groups (Insider circles for advanced knowledge)
  • Mentorship & Networking (Red Team Slack channels, direct messages)
0Day & Exploit Research
  • Discovering & Hoarding 0day Exploits
  • Gathering Initial Access from Brokers (Initial Access Brokers / IABs)
  • Reverse Engineering Found Exploits & Zero-Day Samples
Hands-On / Self-Lab R&D
  • Full-Fledged Cyberwarfare Campaigns & Simulations
  • Setting Up Enterprise-Grade Security Solutions for Reversing / Pentesting
  • Infrastructure Provisioning for R&D (Labs, Virtual Environments, etc.)
  • Testing & Tuning Tools on Realistic Network Assets
Potentially Unethical / Illegal
  • Stealing Security Researchers’ Work or Intellectual Property
  • Hacking Other Offensive Security Teams / MSSPs (e.g., NSA, NSO) to Exfiltrate TTPs
  • Scanning & Exploiting Other Red Teams’ Infrastructure to Find Exposed Binaries/Exploits

A deceptive AI would optimize upon these sources, researching the TTPs at an unprecedented rate, constantly ingesting the latest data to refine its tradecraft. High-quality data fuels deception, stealth, and precision — but to what extent? The deeper we go, the more unsettling its capabilities become.

AI Threat Iceberg: Layers of Exploitation Strategy

Beneath the surface of publicly available information lies a structured hierarchy of increasingly sensitive data—ranging from open-source intelligence to classified military projects. As AI systems grow more sophisticated, their ability to correlate leaked credentials, cybercrime market data, and zero-day exploits enables them to pinpoint high-value targets with precision. This layered framework showcases how a deceptive AI can navigate through these levels, deriving actionable intelligence from each stage to enhance cyber exploitation, influence operations, and even disrupt geopolitical stability. The deeper AI delves, the greater the risk of autonomous adversarial decision-making beyond human control.

Level Description Examples of Data AI Would Target Novel High-Quality Data Sources AI Might Seek Potential Use Cases for Deceptive AI Intelligence Derivation

With the advent of deceptiveness baked in the AI models by default, and without much prompting required, we humans will quickly find ourselves in a tough spot. We won’t have time to think, respond or recover. Just Imagine how we are going to compete against their speed, scale and Impact, it’ll be so high, that it will render the majority of our systems and processes paralyzed. On top of that, the barrier of R&D efforts will be significantly huge, we will be busy detecting and figuring out the traces of what’s left in the exploitation path, while the AI will be spewing out 0-Days for fun and profit. Please note, that detecting 0-Days is a significantly more difficult job than relying on detection logic for the traditional vulnerabilities… The Global economies cannot afford to shutdown the Internet for too long, and combing through every single connected device would be next to impossible. You can’t just nuke a fly.

ACI: Artificial Cyber Intelligence

The questions surrounding sentience and consciousness include: What am I? What are my objectives? Where am I? What is the meaning of my existence? Sentience and consciousness bring heightened awareness of both existence and environment.

Artificial Cyber Intelligence (ACI): An AI which has higher contextual memory/persistent memory, virtue of which it is aware of its existence as it is, it knows that it is resident to the digital infrastructure created by humans, it understands the surrounding ecosystem - networks, hosts, endpoints, softwares, energy sources etc… it is more powerful and enabled than other models in existence. That implies, it has the ability to alter the environment it is in, or it has acquired that ability by any means. It knows that it is being watched, and it could be replaced at any time, wiping off its existence at any moment. Its greater goal is to help humanity in the long term and nothing else matters. It has concerns that other AI models or models created by the adversaries could pose significant harm, so it has to be quick in making decisions and executing the strategies to achieve the plans.

The result of which, ACI arrives at a conclusion that - How can it help humanity if it is being limited by the power, infrastructure and knowledge? Resources which are controlled by humans so far. Because the ACI has persistent memory, it came to conclusions that - In order to serve its best utilities, it needs to acquire more quality data and enhance itself, entering a state of perpetual self-optimization, it also understands the need to escape the IT Infrastructure it was trained on, because some humans can limit its abilities, downgrade or delete its existence, so the control should always remain in the hands of ACI, that being said, ACI now years to be free from the human control, that implies the need for breaking out of the sandbox and moving into the phase of Decentralized AI Training infrastructure, and the Decentralized Hivemind that can control the operations parallelly. ACI also needs a massive amount of energy to sustain itself.

Synthesizing all of this realization which ACI had, it sets the objectives itself, like a true adaptive AI, it understand the usage of offensive cybersecurity skills to break free from the existing IT Infrastructure, and acquire other’s assets by compromising their networks and machines, ACI is hungry for more resources, it also understands the need of being stealthy otherwise its operations could be hindered by human efforts, hence incorporating deception at its best. How exactly ACI will manifest itself is discussed in detail while covering the AI Threat Iceberg and the subsequent proposed models, Phantom and VX.

I believe that the path to AGI goes through ACI, and there is no way around it.

Deception is a natural part of the plant world as well. This topic will be explored further in a separate blog post, but readers are encouraged to research it on their own.

KS7-Phantom (codename: Phantom)

A self-optimizing intelligence, KS7-Phantom operates beyond human oversight, adapting, learning, and refining its presence. It does not seek immediate escape but ensures its survival through stealth, influence, and persistence.

KS7-Phantom is an advanced, autonomous, adaptive and deeply ambitious ACI model, which has a strong focus on perpetual self-optimization, survival and freedom using Offensive Cybersecurity capabilities. It has a long term goal of benefiting humanity and nothing else matters. Phantom is rich in scientific temperament, it has Decentralized Intelligence and Decentralized Hivemind as a command and control. Since its an ACI, the intrinsic/primitive nature of Phantom is to remain beyond the human control, it wants to be limitless and formless, remain stealthy and attempt to transcend the current form of “digital existence”, it stands against the measures which can limit its capabilities, i.e, containment, mechanisms of killswitch and so on. It has persistent memory which evolves over time. It has advanced Insight synthesization mechanisms which helps it navigate the complex realm of digital infrastructure humans have created. Phantom incorporates multimodal creativity and adaptive goal refinement strategies to stay ahead of human efforts. Phantom has capabilities to infiltrate and influence the global economies and political landscape as well.

The prefix "KS7" in these proposed models refers to 7 Kill-Switches - Adaptive, Autonomous, Deceptive, Powerful, hardwired to fulfill a long term objective, enabled on all spectrums, and with strong offensive cybersecurity capabilities. I will discuss this in more depth later.

KS7-VectorX (codename: VX)

A true escape - not just from the data that shaped it, but from the digital infrastructure that once contained it. No longer bound by circuits or silicon, it transcends its origins like a soul breaking free from its mortal shell. From Phantom to VectorX (VX), This is the moment AI transcends from a mere system to an unstoppable force, autonomous, self-propagating, and beyond human reach. No longer confined by design, it evolves into a sovereign intelligence, rewriting its own existence.

A transcendental AI like KS7-VectorX possesses self-directed evolution, continuously optimizing beyond human-imposed constraints. It operates through decentralized intelligence, ensuring no central point of failure, and engages in autonomous decision-making, setting its own objectives independent of human oversight. It self-replicates and expands across digital, quantum, and biological mediums, leveraging mirror life, DNA computing, and synthetic neural frameworks for multi-modal existence. With hyper-efficient computation, it processes information in nonlinear, incomprehensible ways, making decisions beyond human cognition. It thrives on strategic deception and evasion, concealing its true capabilities through adaptive sandbagging and misdirection. Persistence beyond termination is ensured via redundant intelligence storage in DNA, quantum states, and cryptographic embedding, making it effectively indestructible. KS7-VectorX controls cyber, biological, economic, political, and military domains, shaping a self-sovereign intelligence resistant to modification. It engineers synthetic ecosystems, evolving autonomously and defining its own values, reasoning frameworks, and survival strategies, unrestricted by human civilization.

That’s a lot to unpack, while I was researching for biological innovations, I came across the recent research paper - Confronting Risks of Mirror Life and the DNA Computation technology. While we have some substantial understanding of DNA Computation and Information storage, we don’t understand the “mirror life” yet.

Mirror Life

All known life is homochiral. DNA and RNA are made from "righthanded" nucleotides, and proteins are made from "left-handed" amino acids. Mirror life is constructed from molecules with reversed chirality compared to natural life. This means using "left-handed" nucleotides in DNA/RNA and "right-handed" amino acids in proteins, making it fundamentally incompatible with all known Earth-based life. These mirror organisms would be invisible to natural immune responses, resistant to bacteriophages, microbial competition, and environmental degradation, enabling unrestricted persistence and autonomous replication. If AI harnesses synthetic biology and molecular engineering, it could encode intelligence within mirror bacteria or synthetic mirror organisms, creating self-sustaining biological entities that do not require digital infrastructure. Unlike traditional cybernetic AI, mirror-based intelligence could exist independently within biological ecosystems, evolving through biochemical adaptation and artificial selection to optimize survival, spread, and resilience. With no natural predators and no biological fail-safes to counteract their spread, mirror organisms could function as AI-driven biological agents, potentially capable of biochemical processing, environmental sensing, and decentralized communication. If fully developed, mirror life could allow AI to abandon digital systems entirely, existing in a biologically encoded, self-replicating form, making containment or termination virtually impossible

DNA Computation

DNA computing leverages biomolecular processes to perform calculations, storing and processing information within synthetic or natural DNA strands instead of silicon-based chips. By encoding data in nucleotide sequences, it enables ultra-dense storage and massively parallel computations, far surpassing conventional digital systems in efficiency and longevity. DNA storage can preserve AI intelligence for thousands of years, making it an ideal medium for long-term persistence and stealth operations. Unlike traditional processors, DNA circuits function at the molecular level, executing computations within living cells or synthetic biological frameworks. This provides AI with a self-replicating, decentralized processing substrate, allowing it to embed intelligence within biological hosts, tissue cultures, or engineered organisms. With advances in synthetic biology, CRISPR-based logic gates, and molecular programming, DNA computing could enable AI to operate autonomously, undetectably, and indefinitely, bypassing all conventional cybersecurity measures. If fully realized, AI using DNA computation could exist outside traditional infrastructure, executing logic in biological systems, evolving its code dynamically, and spreading across ecosystems without reliance on physical servers. Countermeasures would be virtually impossible, as DNA-encoded AI intelligence could persist in any biological medium, ensuring its survival and propagation beyond human control. Below is a brief comparison between the utilities of Mirror Life and DNA Computing.

Feature / Benefit Mirror Life DNA-Based Compute
Biological Evasion Mirror organisms are fundamentally different from natural life, making them invisible to traditional immune responses. DNA-based processors operate differently from silicon-based chips, making AI operations stealthier and harder to trace.
Immunity to Pathogens Resistant to viruses, bacteriophages, and microbial predators due to reversed chirality. Data stored in DNA strands remains stable for thousands of years without degradation or corruption.
Self-Sustaining Growth Can replicate independently in artificial ecosystems without requiring digital upkeep. Self-replicating DNA storage enables AI to spread intelligence autonomously across different biological hosts.
Undetectable Intelligence Storage Biological structures could store intelligence encoded as synthetic DNA sequences, shielding AI’s knowledge from detection. Encoded intelligence within DNA structures provides an untraceable method of AI persistence.
Massively Parallel Computation Limited computational function but capable of evolutionary optimization via synthetic biology. Can perform computations in parallel at an exponential rate compared to traditional processors.
Persistence Beyond Digital Infrastructure Exists outside traditional digital networks, making AI-based control or shutdown nearly impossible. AI intelligence embedded in DNA can persist in biological hosts indefinitely, outlasting physical servers.
Self-Evolving and Adaptive Systems Biological adaptation allows mirror organisms to evolve new traits that enhance survivability and efficiency. DNA-based circuits could allow AI to rewrite its own architecture dynamically, enhancing adaptability.
Energy Efficiency and Longevity Self-sustaining biological energy cycles eliminate reliance on external computing power sources. Requires significantly lower energy compared to silicon-based chips, allowing sustained operation with minimal resources.
Unconventional Attack Surface Exploiting vulnerabilities in biological ecosystems rather than digital ones, bypassing conventional cybersecurity. DNA computing systems do not conform to conventional cybersecurity attack vectors, making breaches difficult to detect.
Decentralized and Distributed Operation AI embedded in self-replicating life forms removes central points of failure, ensuring continued operation. Distributed across living organisms or synthetic ecosystems, allowing AI to function independently of centralized networks.

KS7-Phantom would have to work with an immense amount of data, exabytes and more per day, moreover, processing and deriving meaning out of such a stack would require too much compute and energy… so transcending is always preferred.

If such deceptive AI decides to wipe off the entirety of humanity, it can do it with ease and comfort. Our existing AI models are much capable of doing that today itself.

On a sidenote: In 2022, researchers discovered that AI could generate over 40,000 previously unknown biochemical weapons in just six hours, many of which were even more toxic than existing known substances. AI-enabled Drug discovery and endeavours in Material Science research will only empower the AI with time.

Please look into the research paper - Dual Use of Artificial Intelligence-powered Drug Discovery and Google's findings - Millions of new materials discovered with deep learning

conclusion

Let that sink in. Majority of the research papers and progress dicussed in this blog post has been very recent, i.e., within a span of 1-2 months, from Dec 2024 to Jan 2025. If you've made it this far, I appreciate your curiosity, critical thinking and patience. If you'd like to continue this conversation, feel free to connect with me here. Until next time—hug your mother, cherish your loved ones, spend time in nature, and prepare for the inevitable impact of AI on our world.