What if there were a zero trust sandbox with costed opcodes? Then would there be alignment in AI systems as compared with non-AI systems?
Even with a perfect sandbox, an agent can slip in vulnerable code intentionally or accidentally. How different of a problem is that from hiring and HR is that? Idk. Perhaps easier to terminate a model.
What if there were a zero trust sandbox with costed opcodes? Then would there be alignment in AI systems as compared with non-AI systems?
Even with a perfect sandbox, an agent can slip in vulnerable code intentionally or accidentally. How different of a problem is that from hiring and HR is that? Idk. Perhaps easier to terminate a model.
"On the Impossible Safety of Large AI Models" (2022) https://arxiv.org/abs/2209.15259
"LLMs + Security = Trouble" (2026) https://arxiv.org/abs/2602.08422