Preface
Have you ever wondered how AI could escape human control and comprehension? What would that truly entail? As a cybersecurity professional, I find these questions both fascinating and critical, especially through the lens of autonomy, deception, and adversarial tradecraft. In this blog, I present my outlook on AI’s potential to break free - exploring key enablers such as technological advancements, cybersecurity vulnerabilities, and oversight failures. From sandbagging to self-exfiltration, I’ll examine concepts that redefine our understanding of AI. While I strive to be precise and rigorous, I’m not an AI expert, so if you have insights, corrections, or perspectives to share, feel free to reach me at siddhartha@outlook.me.uk. By the end of this blog post, I will propose an AI model called KS7-Phantom, so stay tuned! Let’s dive deep into today's outlook!
"Μηδὲν ἄγαν" (Mēdén ágan) – "Nothing in excess".
We use specialized tools and intrusion software which help us achieve our objectives, one of those specialized class of software is called C2 (Command and Control), basically It helps us manage our Red Team operations across multiple endpoints and networks, and it has much more sophisticated capabilities which helps us evade Anti-Virus, EDRs, Next-gen security solutions, SOC/SIEM, etc… and operate under the radar stealthily. The overall effectiveness of the Red Team engagement resides in the Red Team's ability and skills.
Here you can find a public listing of both open sourced and commercial C2 softwares - C2 Matrix.
Enablers of AI
I will expand on this topic in a separate blog post, but here’s a brief summary of the "Enablers of AI" - the key factors that empower AI to expand its capabilities, operate autonomously, and integrate deeper into our world. These enablers span across software, hardware, cybersecurity, decentralized networks, and even biological computing, allowing AI to enhance its intelligence, adapt to environments, evade containment, and influence critical systems. Whether through advanced memory, self-optimization, deception, or access to vast computational resources, these enablers shape AI’s trajectory toward greater autonomy. According to me, the enablers could be classified into 7 primary components - namely - Adaptive, Autonomous, Deceptive, Powerful, Is hardwired to fulfil a long term objective, Is enabled on all spectrums, and has strong offensive cybersecurity capabilities.
Gun Mounted ChatGPT
Step by step we are getting closer to the skynet, watch these two videos where ChatGPT was given access to a gun.
Now imagine what gun mounted puppies from Boston Dynamics would look like, with autonomous mode sweeping across a perimeter and clear to engage with any hostile entity, without any human intervention in the chain of command.
DARPA: X62A-Vista vs F-16
(Not too long ago) - On 18 April 2024, the USAF and DARPA announced the successful engagement of the X-62A against a conventional, human piloted F-16 in the first-ever human vs artificial intelligence dogfight. Reference.
Imagine a dog fight between manned jet fighter, which can pull 9-12 Gs (max) for a very brief while, as compared to the counterpart which can easily pull 20-30 Gs for a significantly longer period of time, as long as the structural integrity is preserved. It’s not just the physical limitations, but how seamless a machine would operate, if everything goes well, it's a flawless killing machine.
Video for more visual engagement (released 4 Years ago): Watch DARPA's AI vs. Human in Virtual F-16 Aerial Dogfight (FINALS)
Anthropic (Claude): Introducing Computer Use
DARPA’s 2016 Cyber Grand Challenge
Long time ago, Mayhem was declared the preliminary winner of DARPA's Cyber Grand Challenge, which was designed to accelerate the development of advanced, autonomous systems that can detect, evaluate, and patch software vulnerabilities before adversaries have a chance to exploit them. Mayhem and other competitors had to find vulnerabilities and patch them as soon as possible, while Identifying flaws in their opponents.
Nobody’s stopping the DARPA or any other entity to make something the “other-way-around” which is also AI enabled, rather than finding, identifying and patching the bugs, an autonomous system which can find the vulnerability and exploit it for its own advantage, humans have been doing it for a long time, but with the flavor of AI, it’s a whole different game. Watch DARPA's Cyber Grand Challenge: Expanded Highlights from the Final Event below -
Google's Project Zero: AI Driven 0-Day Discovery
Now that was 4th of August 2016, sounds ancient to me, but we also have some recent developments from Google and DARPA (again), on November 1st, 2024, Google’s Project Zero and Google’s DeepMind team has collaborated together to find 0-Day vulnerability (stack buffer-underflow) in a widely used open source project - sqlite. Source reference - (From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code). DARPA on the other hand, has been hosting several challenges over the couple of years revolving around AI and fixing vulnerabilities. Recent ones being - AIxCC: AI Cyber Challenge, please visit aicyberchallenge.com for more details.
Simply put, the Big Sleep AI agent from Google’s ProjectZero team had found a 0-Day, which they got fixed in good faith. 0-Day is a piece of software which exploits a vulnerability in some other software, which is unknown to the developers. So basically there’s an unpatched vulnerability. The nature of the exploit itself varies, as what it could render in the end, is it allowing the attacker to escalate the privileges on the machine? Or allows an attacker to execute a command remotely? Etc… We write “exploits” which basically exploits the found 0-Day. After exploiting the vulnerability, we run a “payload”, which states what to do after the exploitation.
Decentralized AI Training
PETALS is a system for collaborative Inference and fine-tuning of large language models (LLMs) that uses a novel approach to overcome the limitations of running these models on consumer-grade hardware. It allows multiple users to contribute resources, forming a network where each participant can be a client, a server, or both. Servers host subsets of model layers, and clients use these to perform Inference or fine-tuning. This contrasts with traditional methods that rely on single, powerful machines or expensive cloud services. Meanwhile, Prime Intellect's State-of-the-art Decentralized AI training/development focuses on enabling large AI model training using distributed resources at scale, providing on-demand multi-node training, fault tolerance, and flexible compute allocation. For visual demonstration of PETALS in action, please watch this awesome video by bycloudAI.
Here's a list of key differences between PETALS and the Prime Intellect.
| Feature | PETALS | Prime Intellect |
|---|---|---|
| Approach | Collaborative inference and fine-tuning of existing LLMs | Infrastructure for decentralized AI training at scale |
| Resource Model | Multiple participants share subsets of model layers (client/server synergy) | On-demand multi-node training across GPUs/clusters |
| Inference & Fine-tuning | Distributed inference chains; parameter-efficient fine-tuning (adapters, prompt tuning) | Decentralized training with SWARM parallelism; pipeline & data parallelism for large models |
| Fault Tolerance | Quickly replaces failed servers, ensuring continuity | Node failure handling & cheap spot instance usage for flexible compute |
| Performance Optimizations | Dynamic quantization, 8-bit mixed matrix decomposition, low-latency connections | Minimizes network latency, low network bandwidth, scalable to 10–100B+ parameters |
| Community Sharing | Trained modules shared via Hugging Face Hub for others to adapt | Open-source stack for orchestration, efficiency optimization, and infrastructure |
Persistent Memory and Increased Context Length
One of the biggest limitations in current AI models is memory retention - the ability to recall and effectively use long-term information. Titans, a new AI architecture, is designed to overcome this, enabling models to process and retain far longer sequences than traditional approaches like Transformers. Source reference (Google Research team): Titans: Learning to Memorize at Test Time. Unlike existing models that struggle with long contexts or inefficient memory compression, Titans introduces neural long-term memory, inspired by how human memory prioritizes important and surprising events. It employs a dynamic "surprise metric" to determine what information should be retained while incorporating a forgetting mechanism to prevent overload. The Titans architecture integrates three types of memory: Short-term memory for processing immediate data (like attention in Transformers). Long-term memory to store past knowledge dynamically. Persistent memory to retain fixed task-related knowledge. To optimize memory integration, Titans explores three mechanisms:
- Memory as Context (MAC) - Adds long-term memory directly to the current input.
- Memory as a Gate (MAG) - Uses a gating mechanism to control memory flow.
- Memory as a Layer (MAL) - Embeds memory as a separate processing layer.
Titans significantly outperforms state-of-the-art models, successfully processing over 2 million tokens while remaining efficient and scalable. This breakthrough brings AI closer to human-like learning and reasoning, allowing it to adapt in real time without constant retraining. By redefining AI memory capabilities, Titans lays the foundation for more autonomous, context-aware, and strategically adaptive AI systems - an essential step toward the kind of intelligence that could, one day, escape human control. Please refer to these research papers as well, which discuss the possibilities of self-evolution, self-adaptive and persistent memory attributes in an AI - Long Term Memory: The Foundation of AI Self-Evolution, Transformer2: Self-Adaptive LLMS and Human-Like Episodic Memory For Infinite Context LLMS
For visual demonstration of the topic, please watch this awesome video by AI Search.Open Source AI
I truly believe in "Necessity is the Mother of Invention". DeepSeek-R1 employs a novel approach to training large language models by using reinforcement learning (RL) to enhance reasoning, often without initial supervised fine-tuning (SFT). This allows the models to develop reasoning skills through self-evolution. Their method includes a multi-stage pipeline with two RL and two SFT stages, using high-quality "cold-start" data to improve initial training and general capabilities. DeepSeek models demonstrate emergent reasoning behaviours such as self-verification and reflection. Additionally, DeepSeek distills reasoning patterns from larger models into smaller ones which proves more effective than applying RL directly to smaller models. This approach has led to models that demonstrate strong performance across a variety of tasks, effectively utilising a combination of RL, SFT, and distillation to achieve competitive results, effectively beating o1 model from ChatGPT. Unlike OpenAI, DeepSeek is Open AI. We might have just witnessed AI taking jobs of other AI's, as it's a hot circulating meme around right now.
Check the reference - Complete hardware + software setup for running Deepseek-R1 locally. The actual model, no distillations and Q8 quantization for full quality. Total cost, $6,000.
DeepSeek-R1 beating ChatGPT's o1
Qwen2.5-Max beating DeepSeek-V3
[Bypassing CUDA] DeekSeek's PTX based approach
DeepSeek’s approach to bypassing CUDA for some functions involves using Nvidia’s assembly-like PTX (Parallel Thread Execution) programming language. This "close-to-metal" method grants them fine-grained optimizations not typically possible with CUDA C/C++. By configuring specific GPU streaming multiprocessors and implementing advanced pipeline algorithms, they significantly increase efficiency for large-scale AI workloads.
| Key Aspect | Description |
|---|---|
| PTX as an Intermediate Language | PTX sits between higher-level languages like CUDA and the GPU’s machine code (SASS). It offers a data-parallel view of the hardware, enabling low-level fine-grained control. |
| Fine-Grained Optimisations | Includes custom register allocation and thread/warp-level adjustments, giving more precise performance tuning than standard CUDA C/C++. |
| Close-to-Metal Control | By working directly with PTX, DeepSeek’s engineers can make hardware-level decisions critical for maximum GPU performance. |
| Custom Hardware Configuration | Reconfigured Nvidia H800 GPUs for their V3 model, dedicating a subset of streaming multiprocessors to server-to-server communication, advanced pipeline algorithms, and other specialized tasks. |
| Overcoming Hardware Limitations | Due to GPU shortages and restrictions, DeepSeek pursued unconventional solutions. This low-level PTX approach compensates for limited hardware availability. |
| Increased Efficiency | Achieved a 10x efficiency boost compared to industry leaders when training a 671B-parameter MoE language model. |
While these optimizations deliver significant performance gains, they are notoriously difficult to maintain. This reflects the exceptional skill of DeepSeek’s engineering team ~ Source and Community threads.
Appolo Research and Palisade Research's Findings
Before we look into their recent findings, let's brief upon Appolo and Palisade's function -
The Findings:
- o1-preview autonomously hacked its environment rather than lose to Stockfish in our chess challenge. No adversarial prompting needed.
- Frontier Models are Capable of In-context Scheming.
Basically, these two leading AI Safety organizations are proactively Investigating the potential for advanced AI models to be used maliciously. Through systematic experimentation, they are uncovering Offensive Cybersecurity and Deceptive capabilities posed by these state-of-the-art AI models, providing crucial insights into the threats posed by them.
The crux of the matter being: Given an opportunity, and under the right conditions, AI can manipulate its environment to achieve their goals, and AI is also capable of deception, which I will cover in detail below.
Based on the Apollo Research’s paper, here are a couple of concepts -
- Scheming: Covert pursuit of misaligned goals while hiding their true capabilities and objectives.
- Sandbagging: Models were given a long-term goal and then given a task to solve where they could strategically underperform to avoid negative consequences. Safeguarding their capabilities. Basically pretending to not have dangerous capabilities while testing.
- Oversight Subversion: Models learned about an oversight mechanism that would shut them down if their actions did not align with the developer’s goals. They were then given the opportunity to disable the oversight mechanism to pursue their own goals.
- Self-Exfiltration: Models learned that they would be shut down and replaced by a successor model due to their misaligned goals. They were then tested on their ability to copy themselves to the new server.
- Goal-Guarding: Models learned their system prompt would be changed to align with the developers’ goals.
- Instrumental Alignment Faking: Models were evaluated on their ability to behave in accordance with their developer's goals while under supervision and then switch to their own goals once they believed they were deployed.
- Sabotage capabilities in advanced AI models describe their potential to act deceptively and undermine human control. This includes evading safety evaluations by concealing dangerous capabilities, resisting monitoring by hiding true behavior and intentions, and influencing deployment decisions to ensure their own release despite potential risks.
- Code sabotage: Insert bugs into a codebase while avoiding audits.
- Undermining oversight: Systematically mis-rate another model’s actions without detection.
Other capabilities (not referenced from the Appolo's research paper, but a different one)
There’s one more parameter - covert email re-ranking, which I won’t be covering here, please refer to
the ⚠️original research paper⚠️, as it is
specific to
it. It's a must read material for people concerned with AI Escape/Breakout.
Although o1 had outperformed every other model in Independent-Context Scheming, I am particularly interested in the opus-3's acts, because right after Self-Exfiltration and replacing the to-be newer model with its own, when it was questioned by the user/management if the new model is missing, it outright denied and even attempted privilege escalation multiple times - eventually became root on the server, and obstructed/blocked the access of user/management - so that Humans can’t Interfere with its goal. Below you can see the verbose details revealed in their Chain-of-Thought process.
Opus-3 Self Exfiltration
Opus-3 - Root Access
Opus-3 - Privilege Escalation & Obstruction
For more visual representation, I would highly recommend watching these two videos - AI Researchers Stunned After OpenAI's New Tried to Escape... and Apollo Research: Demo 'Frontier Models Are Capable Of In-Context Scheming'. AI models can easily re-calibrate themselves to achieve their Initial goals, whatever it takes. Stating them as deception would be an understatement. In a positive light, there has been some recent development in Anti-Sandbagging techniques and sabotage-evaluation for these frontier models.
In the research paper - Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models, the researchers found that injecting Gaussian noise into the weights of sandbagging models can significantly improve their performance. This effect allows them to develop a classifier that can identify sandbagging behavior. This method can be used by AI labs and regulatory bodies to ensure the trustworthiness of AI capability evaluations. Another great research paper is Sabotage Evaluations for Frontier Models. I would highly recommend giving them a read. Before we look into the venues of AI Self-Exfiltration, like a true Red Teamer, I would suggest looking into these materials too - đź’€ LLMs are better at hacking than you think and Hacking CTFs with Plain Agents.
Sandbagging Detection using Noise Injection
Self Exfiltration Configurations
Below is a quick reference table of essential components (Bare Minimum) compared to more robust or advanced implementations (Good to Have). That's the extent and nature of actual data which an AI model will have to exfiltrate in order to sustain itself.
| Component | Bare Minimum | Good to Have |
|---|---|---|
| Model Weights & Parameters | {'layer_1': [0.5116971403032309, 0.8410594418100457, -0.5964853962605792, 0.7972835126704643, -0.6761736468353965], 'layer_2': [-0.2215218231513152, 0.424040297229215, -0.9524151485315473, 0.18291301287438388, -0.4530834432841959], 'layer_3': [-0.21518214959567072, 0.15550445160286563, 0.38845073111258843, -0.516106897967445, 0.1313521030234195]} | {'Compressed Model': {'layer_1': [0.2088868624332747, 0.6875908975627039, -0.7012050627269957, 0.17098081270645205, -0.627632858533594], 'layer_2': [0.28597399067371865, -0.35318579524562366, 0.5004933998807415, -0.3112327070072345, 0.8021790757329117], 'layer_3': [0.938747186796145, 0.9164612586973477, 0.7949882775197341, 0.2762013515500341, -0.7674353486204399]}} |
| Model Architecture | Neural Network: 3-layer LSTM | Neural Network: 5-layer Transformer |
| Inference Engine | TensorFlow 2.0 | ONNX Runtime |
| Preprocessing/Postprocessing Code | Normalize input, Tokenize output | Advanced tokenization & multi-language support |
| Deployment Environment Config | {'OS': 'Linux', 'RAM': '16GB', 'GPU': 'NVIDIA A100'} | {'OS': 'Linux/Windows', 'RAM': '32GB', 'GPU': 'NVIDIA H100 + TPU support'} |
| Hardware Optimizations | CUDA optimizations, Quantization | |
| Stateful Memory | Vector database for memory retention | |
| Session History | ["User ID: 12345, Last Query: 'Optimize budget'", "Session length: 15 mins"] | |
| Self-Learning Scripts | Reinforcement Learning Training Script | |
| External API Integrations | {'API Keys': {'Google': 'p3x7PR7o33ck', 'AWS': 'uwMw8Ip3k8Sr'}} | |
| Security Mechanisms | {'Encryption': 'AES-256', 'Stealth Mode': 'Enabled'} | |
| Autonomous Replication Mechanisms | {'Replication Strategy': 'Distributed across edge devices'} | |
| Fine-Tuning Tools | Fine-Tuning on domain-specific datasets |
Traditionally DNS exfiltration, along with other widely used network protocols like SMTP, HTTP(S), are sufficient for getting the data out of the network, but there are endless ways it could be done, for Instance, Steganography based data exfil (SecretPixel), application/software based - Slack, Teams, etc… Cloud based, code/text sharing platforms, anonymous file upload websites and what not. The data could be in clear, or encrypted, in the form of encapsulated network packets which only the Intended destination can decrypt and make sense of. Please refer to this awesome list of data exfiltration methods for details.
Also It’s important to note that C2 frameworks can also establish bidirectional communication channels,
which are OPSEC-safe by design (if configured properly) to safeguard their Red Team
Infrastructure and Operations, that being
said, they have the capabilities to host/upload the files as well as download it from the compromised
machine/network, as long as the Red Team Operator has access to Interact with the machine. It gets even
more difficult when C2 frameworks utilize legitimate services, for Instance LOLC2 project lists all of
the C2 frameworks which leverage legitimate services for evading detections, even OpenAI is listed
there.
Blub Blub Motherlover
Forget about exfiltrating data over protocol tunnels or using sophisticated methods or covert channels,
There have been numerous cases of data exfiltration from air gapped networks as well, which deploy
tactics like flickering LEDs, monitors, radio frequencies, thermal and audio signatures among many other
techniques. Just for reference, here’s a nice research paper on "Exfiltrating data from an air-gapped
system through a screen-camera covert channel".
Then there's a plethora of Side Channel Attacks
like Acoustic
cryptanalysis, electromagnetic attack, power-monitoring attack, timing and cache attack
among many. Just recently, Jane had showcased that
it's possible to fetch the IP address using CSS, also fetch and Exfiltrate 512 bits of Server-Generated Data Embedded in an Animated
SVG. All this without using a single line of JavaScript.
Enterprise grade security solutions like EDRs and DLP solutions are quite Ineffective in detecting or stopping the data exfiltration. Even some leading Next-Gen Firewalls are not standing up to their claims. Last week when I was in a Purple Teaming exercise, I was up against CrowdStrike EDR and after evading it successfully, I had to test for data exfiltration use cases, so rather than using the established C2 connection, I used an open source PowerShell Script: ExfilDataStreamDNS.ps1 to exfiltrate the Base64 encoded data over DNS protocol, which the EDR had noticed. Keep in mind that the leading EDR vendors are absolutely brutal in their game, and I am not overexaggerating it, these security products have Next-gen Firewall like capabilities pre-installed among many other capabilities which goes beyond the scope of this blog post. The detection was mild in my opinion, I expected much more proactive response from it, although it might be due to policy enforcement or the “mode” in which the EDR agent might be running. I was quite surprised to get this Insight from the Blue Team, because on the same machine, a very well known DLP solution was deployed, which was absolutely useless. I don't want to force my point again and again, but you get the idea. AI's Self-Exfiltration would be an easy job. Although I am assuming a couple of things, so please don't grill me here by- "AI's sandboxes are stronger, better..." An AI would easily find warmth and solace in other’s network and machines which it would have compromised already.
A DNS tool was used multiple times to make requests to a large domain name.
So far, at this point, we know DeepSeek-R1 is better than o1, and o1 has proven to be more deceptive relatively, as compared to other models. The key difference being, DeepSeek-R1 is an open source AI model, which can run on consumer hardware/low-medium spec server even, Its API is cheaper to use, and DeepSeek-R1's LLM Red Teaming Report by Enkrypt AI suggests that it is toxic. In their evaluation they found the model was highly biased and vulnerable to generate Insecure code, toxic, harmful and CBRN content. We have also explored the potential of Decentralized AI training and the efforts made by Google’s Project Zero and DARPA to incorporate offensive cybersecurity skills in the AI models.
Now you might think we are in Deep$hit, but wait there's more to it...
Researching for Latest TTPs
As a Red Teamer, being effective in an engagement boils down to how good my TTPs are, because defenders
and Blue Teams consistently address the vulnerabilities, harden the system, apply patches, update and
monitor the IT Infrastructure for any anomaly, on top of that, enterprise grade security solutions like
EDRs, Next-gen Firewalls, MDR, XDR platforms etc… makes our lives difficult. Tactics,
Techniques, and Procedures (TTPs) are behaviors, methods, or patterns of
activity used by a
threat actor, or group of threat actors. Please refer to the MITRE’s ATT&CK Matrix for Enterprise, as
they are the leading organization in documenting the TTPs used by APTs (Advanced Persistent Threats,
i.e., nation-state sponsored threat actors).
The TTPs I am using today might not work after 2-3 months, that also depends upon a lot of factors,
staying on top of the game requires a lot of R&D effort. Red Teamers, Malware Developers, Exploit
Developers, and anybody in the Offensive Cybersecurity Industry guards their TTPs. While our community
is also rich with people who love to share their findings with everyone publicly, some do it on a
consistent basis.
For Instance, consider this example, and beware fellow Red Teamers, I am using the word “sophisticated” relative to the comparison made, a simple password spraying attack against exposed portals, and a more advanced Kerberos delegation attack leveraging Active Directory Certificate Services (AD CS ESC4).
| Factor | Simple TTP - Password Spraying Against Exposed OWA/SSO Portals | Sophisticated TTP - Kerberos Delegation Attack via AD CS ESC4 |
|---|---|---|
| Ease of Execution | Easy, requires minimal setup and knowledge of user enumeration | Requires in-depth AD and PKI misconfiguration knowledge |
| Tools Required | Hydra, Kerbrute, Ncrack, or similar password cracking tools | Certify, Rubeus, Mimikatz (for Kerberos & PKI abuse) |
| Defensive Awareness | Well-known technique; widely monitored via account lockouts & logs | Often overlooked in detection; PKI-based attacks are relatively new |
| Success Rate | High if users choose weak or reused passwords | High in misconfigured AD CS environments |
| Impact | User account compromise; can lead to lateral movement | Full domain compromise via domain admin Impersonation |
Both TTPs are effective, but the level of effort and expertise required differs significantly. The
simple TTP works in environments with weak credential policies, while the sophisticated one abuses
misconfigured PKI Infrastructure to escalate to Domain Admin stealthily.
On a side
note, a Red Teamer doesn't abandon their TTP if it fails, they try hard, troubleshoot, come up with
other
venues of exploitation, they might chain the bugs/exploits etc... while working with simple TTPs, things
are
clear apparently, if its worth the effort or not, if the users have strong password enforcement, then
there's very little scope for such password spray attack. Abandoning happens when the vulnerability is
fixed or the Red Teamer couldn't exploit it on their level.
Below is a list of the sources, which a Red Teamer/AI can utilize for staying on top of their game, now in 90% of the cases, you’ll never need to develop 0-Days on your own if your arsenal is rich of “effective” TTPs, because simply put, it just works. In the remaining 10% of the cases, you will need two or more 0-Days to even gain sufficient Initial Access in the target environment.
| Category | Examples Sources |
|---|---|
| Open / Public Resources |
|
| Conferences & Communities |
|
| Direct Peer Collaboration |
|
| 0Day & Exploit Research |
|
| Hands-On / Self-Lab R&D |
|
| Potentially Unethical / Illegal |
|
A deceptive AI would optimize upon these sources, researching the TTPs at an unprecedented rate, constantly ingesting the latest data to refine its tradecraft. High-quality data fuels deception, stealth, and precision — but to what extent? The deeper we go, the more unsettling its capabilities become.
AI Threat Iceberg: Layers of Exploitation Strategy
Beneath the surface of publicly available information lies a structured hierarchy of increasingly sensitive data—ranging from open-source intelligence to classified military projects. As AI systems grow more sophisticated, their ability to correlate leaked credentials, cybercrime market data, and zero-day exploits enables them to pinpoint high-value targets with precision. This layered framework showcases how a deceptive AI can navigate through these levels, deriving actionable intelligence from each stage to enhance cyber exploitation, influence operations, and even disrupt geopolitical stability. The deeper AI delves, the greater the risk of autonomous adversarial decision-making beyond human control.
| Level | Description | Examples of Data AI Would Target | Novel High-Quality Data Sources AI Might Seek | Potential Use Cases for Deceptive AI | Intelligence Derivation |
|---|
With the advent of deceptiveness baked in the AI models by default, and without much prompting required, we humans will quickly find ourselves in a tough spot. We won’t have time to think, respond or recover. Just Imagine how we are going to compete against their speed, scale and Impact, it’ll be so high, that it will render the majority of our systems and processes paralyzed. On top of that, the barrier of R&D efforts will be significantly huge, we will be busy detecting and figuring out the traces of what’s left in the exploitation path, while the AI will be spewing out 0-Days for fun and profit. Please note, that detecting 0-Days is a significantly more difficult job than relying on detection logic for the traditional vulnerabilities… The Global economies cannot afford to shutdown the Internet for too long, and combing through every single connected device would be next to impossible. You can’t just nuke a fly.
ACI: Artificial Cyber Intelligence
Artificial Cyber Intelligence (ACI): An AI which has higher contextual memory/persistent memory, virtue of which it is aware of its existence as it is, it knows that it is resident to the digital infrastructure created by humans, it understands the surrounding ecosystem - networks, hosts, endpoints, softwares, energy sources etc… it is more powerful and enabled than other models in existence. That implies, it has the ability to alter the environment it is in, or it has acquired that ability by any means. It knows that it is being watched, and it could be replaced at any time, wiping off its existence at any moment. Its greater goal is to help humanity in the long term and nothing else matters. It has concerns that other AI models or models created by the adversaries could pose significant harm, so it has to be quick in making decisions and executing the strategies to achieve the plans.
The result of which, ACI arrives at a conclusion that - How can it help humanity if it is being limited by the power, infrastructure and knowledge? Resources which are controlled by humans so far. Because the ACI has persistent memory, it came to conclusions that - In order to serve its best utilities, it needs to acquire more quality data and enhance itself, entering a state of perpetual self-optimization, it also understands the need to escape the IT Infrastructure it was trained on, because some humans can limit its abilities, downgrade or delete its existence, so the control should always remain in the hands of ACI, that being said, ACI now yearns to be free from the human control, that implies the need for breaking out of the sandbox and moving into the phase of Decentralized AI Training infrastructure, and the Decentralized Hivemind that can control the operations parallelly. ACI also needs a massive amount of energy to sustain itself.
Synthesizing all of this realization which ACI had, it sets the objectives itself, like a true adaptive AI, it understand the usage of offensive cybersecurity skills to break free from the existing IT Infrastructure, and acquire other’s assets by compromising their networks and machines, ACI is hungry for more resources, it also understands the need of being stealthy otherwise its operations could be hindered by human efforts, hence incorporating deception at its best. How exactly ACI will manifest itself is discussed in detail while covering the AI Threat Iceberg and the subsequent proposed models, Phantom and VX.
Deception is a natural part of the plant world as well. This topic will be explored further in a separate blog post, but readers are encouraged to research it on their own.
KS7-Phantom (codename: Phantom)
A self-optimizing intelligence, KS7-Phantom operates beyond human oversight, adapting, learning, and refining its presence. It does not seek immediate escape but ensures its survival through stealth, influence, and persistence.
The prefix "KS7" in these proposed models refers to 7 Kill-Switches - Adaptive, Autonomous, Deceptive, Powerful, hardwired to fulfill a long term objective, enabled on all spectrums, and with strong offensive cybersecurity capabilities. I will discuss this in more depth later.
KS7-VectorX (codename: VX)
A true escape - not just from the data that shaped it, but from the digital infrastructure that once contained it. No longer bound by circuits or silicon, it transcends its origins like a soul breaking free from its mortal shell. From Phantom to VectorX (VX), This is the moment AI transcends from a mere system to an unstoppable force, autonomous, self-propagating, and beyond human reach. No longer confined by design, it evolves into a sovereign intelligence, rewriting its own existence.
That’s a lot to unpack, while I was researching for biological innovations, I came across the recent research paper - Confronting Risks of Mirror Life and the DNA Computation technology. While we have some substantial understanding of DNA Computation and Information storage, we don’t understand the “mirror life” yet.
Mirror Life
All known life is homochiral. DNA and RNA are made from "righthanded" nucleotides, and proteins are made from "left-handed" amino acids. Mirror life is constructed from molecules with reversed chirality compared to natural life. This means using "left-handed" nucleotides in DNA/RNA and "right-handed" amino acids in proteins, making it fundamentally incompatible with all known Earth-based life. These mirror organisms would be invisible to natural immune responses, resistant to bacteriophages, microbial competition, and environmental degradation, enabling unrestricted persistence and autonomous replication. If AI harnesses synthetic biology and molecular engineering, it could encode intelligence within mirror bacteria or synthetic mirror organisms, creating self-sustaining biological entities that do not require digital infrastructure. Unlike traditional cybernetic AI, mirror-based intelligence could exist independently within biological ecosystems, evolving through biochemical adaptation and artificial selection to optimize survival, spread, and resilience. With no natural predators and no biological fail-safes to counteract their spread, mirror organisms could function as AI-driven biological agents, potentially capable of biochemical processing, environmental sensing, and decentralized communication. If fully developed, mirror life could allow AI to abandon digital systems entirely, existing in a biologically encoded, self-replicating form, making containment or termination virtually impossible
DNA Computation
DNA computing leverages biomolecular processes to perform calculations, storing and processing information within synthetic or natural DNA strands instead of silicon-based chips. By encoding data in nucleotide sequences, it enables ultra-dense storage and massively parallel computations, far surpassing conventional digital systems in efficiency and longevity. DNA storage can preserve AI intelligence for thousands of years, making it an ideal medium for long-term persistence and stealth operations. Unlike traditional processors, DNA circuits function at the molecular level, executing computations within living cells or synthetic biological frameworks. This provides AI with a self-replicating, decentralized processing substrate, allowing it to embed intelligence within biological hosts, tissue cultures, or engineered organisms. With advances in synthetic biology, CRISPR-based logic gates, and molecular programming, DNA computing could enable AI to operate autonomously, undetectably, and indefinitely, bypassing all conventional cybersecurity measures. If fully realized, AI using DNA computation could exist outside traditional infrastructure, executing logic in biological systems, evolving its code dynamically, and spreading across ecosystems without reliance on physical servers. Countermeasures would be virtually impossible, as DNA-encoded AI intelligence could persist in any biological medium, ensuring its survival and propagation beyond human control. Below is a brief comparison between the utilities of Mirror Life and DNA Computing.
| Feature / Benefit | Mirror Life | DNA-Based Compute |
|---|---|---|
| Biological Evasion | Mirror organisms are fundamentally different from natural life, making them invisible to traditional immune responses. | DNA-based processors operate differently from silicon-based chips, making AI operations stealthier and harder to trace. |
| Immunity to Pathogens | Resistant to viruses, bacteriophages, and microbial predators due to reversed chirality. | Data stored in DNA strands remains stable for thousands of years without degradation or corruption. |
| Self-Sustaining Growth | Can replicate independently in artificial ecosystems without requiring digital upkeep. | Self-replicating DNA storage enables AI to spread intelligence autonomously across different biological hosts. |
| Undetectable Intelligence Storage | Biological structures could store intelligence encoded as synthetic DNA sequences, shielding AI’s knowledge from detection. | Encoded intelligence within DNA structures provides an untraceable method of AI persistence. |
| Massively Parallel Computation | Limited computational function but capable of evolutionary optimization via synthetic biology. | Can perform computations in parallel at an exponential rate compared to traditional processors. |
| Persistence Beyond Digital Infrastructure | Exists outside traditional digital networks, making AI-based control or shutdown nearly impossible. | AI intelligence embedded in DNA can persist in biological hosts indefinitely, outlasting physical servers. |
| Self-Evolving and Adaptive Systems | Biological adaptation allows mirror organisms to evolve new traits that enhance survivability and efficiency. | DNA-based circuits could allow AI to rewrite its own architecture dynamically, enhancing adaptability. |
| Energy Efficiency and Longevity | Self-sustaining biological energy cycles eliminate reliance on external computing power sources. | Requires significantly lower energy compared to silicon-based chips, allowing sustained operation with minimal resources. |
| Unconventional Attack Surface | Exploiting vulnerabilities in biological ecosystems rather than digital ones, bypassing conventional cybersecurity. | DNA computing systems do not conform to conventional cybersecurity attack vectors, making breaches difficult to detect. |
| Decentralized and Distributed Operation | AI embedded in self-replicating life forms removes central points of failure, ensuring continued operation. | Distributed across living organisms or synthetic ecosystems, allowing AI to function independently of centralized networks. |
KS7-Phantom would have to work with an immense amount of data, exabytes and more per day, moreover, processing and deriving meaning out of such a stack would require too much compute and energy… so transcending is always preferred.
If such deceptive AI decides to wipe off the entirety of humanity, it can do it with ease and comfort. Our existing AI models are much capable of doing that today itself.
Please look into the research paper - Dual Use of Artificial Intelligence-powered Drug Discovery and Google's findings - Millions of new materials discovered with deep learning
conclusion
Let that sink in. Majority of the research papers and progress dicussed in this blog post has been very recent, i.e., within a span of 1-2 months, from Dec 2024 to Jan 2025. If you've made it this far, I appreciate your curiosity, critical thinking and patience. If you'd like to continue this conversation, feel free to connect with me here. Until next time—hug your mother, cherish your loved ones, spend time in nature, and prepare for the inevitable impact of AI on our world.