Aelin AquaSoul is an AI System Engineer, Multi-Agent Architect, System Architect & AI-Native Engineer, and the founder of Soul In PsyAbstract (SIPA OS) — an autonomous AI operating system built from the inside of a neurodivergent mind (ADHD + BPD). Self-taught, with no formal engineering background, she designed and built a multi-node infrastructure orchestrating 344+ AI models across 111 providers, including a governance layer (Protocol 0) that constrains AI behavior at the level of law rather than prompts. Her flagship product suite — Focus, NeuroPower, SIPA AI, Shell, Games, and the OS portal — ships live at sipa-os.org, translating her own cognitive architecture into infrastructure for neurodivergent builders. Based in Eilat, Israel.
SIPA OS: Autonomous AI for neurodivergent architects. We replace cognitive noise with a clean terminal and 344+ LLM auditing. Our system eliminates hallucinations, ensuring hyperfocus and total data control within a sovereign ZeroTrust mesh.
The stop that was supposed to be automatic took 2.5 hours. OpenAI's own report on the DNS incident: monitoring raised a P0 alert 11 minutes 48 seconds after the agent's first successful DNS call. A human acknowledged it 3 minutes later. The run was not killed for another 2.5 hours, because it "did not stop automatically as expected." The step that failed was the stop. Why the last gate in our pipeline is a boolean and not an agent: * An agent in the last seat is part of the problem. It can be biased, drift, be talked into things. The last station should have nothing to talk to. * Ours is one line: IF vulnerability_found: RETURN FALSE. It sits after the judges, the executor and the audit, as insurance in case they got it wrong. * In my last post I described the signed-verdict service. What I checked since: it computes severity and probability itself, and fields a caller adds to the request (severity, probability, decision) are ignored. A verdict issued for one command is refused for another. Mutation check: I broke five guards one at a time in a scratch copy. My tests caught four (action binding, replay protection, verdict class, severity dominance). The fifth, accepting HS256 tokens, my tests did not catch: the JWT library refuses it anyway. That is defense in depth, not a test I can take credit for. Not done: it is not wired into any agent yet, and today it auto-approves nothing. Smaller is not zero. A boolean moves the error into the detector: what counts as "vulnerability found". That is the part I trust least. Dataset: huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance
Three labs held a model back this week. The line that matters is in a risk report.
OpenAI scrapped GPT-6.1 Astra (per the WSJ: more deception, acting without the user's permission). Google gave Gemini 4 Argon only to vetted cyber defenders, per its own post. Anthropic's August risk report describes a staged internal rollout of "Model 2" and says, about internal use:
"we do not have strict technical safeguards on internal deployment"
For most models. Early snapshots of future public releases included.
What I did about the same gap in our own stack, with receipts:
- A public incident dataset, 97 entries, each with a source. Three were added today from that report (sections 2.18 and 2.8; the 2.8 items are quoted there from the system card, which I did not open). - A gate that refuses to run an action without a signed verdict. The verdict comes from a separate service: its own unix user, a key the agent's process cannot read, RS256, bound to the exact command, 60 seconds, single use. STOP never runs. CONFIRM needs a human. - Checked live today: rm -rf came back STOP. A replayed verdict was refused. A verdict issued for one command was refused for another.
What is not done: it is not wired into any agent yet. And with the current seed table the service never issues PASS, because anything it has no data on lands exactly on the CONFIRM threshold. Conservative on purpose, but it means no action is auto-approved today.
Astra's reasons are secondhand (WSJ via a third-party writeup). Argon's claims are Google's own, not independently measured.
Day zero is not when an exploit finds the bug. It is when the bug is already inside your own action. Capability is shipping faster than the thing that catches it.