CypherCon 2025
Is AI Safety an illusion?
Roman Eng
Abstract:
As an alignment researcher in Artificial Intelligence, I often wonder if we are chasing the wrong safety outcomes or trying to answer the wrong questions. The digital world is expanding exponentially, yet we often assume we can control risks in this virtual world.
We have been spoiled by traditional software applications, which are innately deterministic and have minimal subjectivity. Modern AI language models defy this confidence. They are unwieldy, lack the ground truth we long for, and resist our so-called AI safeguards. Indeterministic by nature, they suffer from ambiguity to no end. If AI safety relies on predictability, how do we align them with humans and real-world use cases that are anything but predictable?
In this presentation, we’ll talk about my conjecture on why no single cybersecurity solution has emerged. We’ll explore why jailbreaking AI systems remains surprisingly straightforward. We’ll then discuss early mitigation techniques aimed at preventing these jailbreaks—highlighting both their promise and their limitations.
As AI continues to scale in capability, we stand at the crossroads because AI safety is not as binary as we hoped. If AI safety is an illusion, then human agency must be our anchor for a positive future. The challenge ahead should not be solely about mitigating risks but about redefining our relationship with intelligent systems. We must ensure they reflect our values and not just the pursuits of power and capability.
Perhaps the real question is not how to make AI safe, but how we redefine our coexistence with intelligence beyond our own.

Roman Eng
Roman Eng is likely AI
Roman Eng is an advisor and technologist with deep expertise in cybersecurity and artificial intelligence. Over 30 years of professional experience in incident response, risk and compliance, and building secure data centers, servicing the largest organizations in the Midwest. Doctoral student in Computer Science at Clarkson University with a research focus in AI Safety, Governance and Explainability. His research aims to contribute to the safety and trustworthiness of machine learning models, preserving humanity’s wellbeing in the era of AI advancement. The “good” in humanity must be preserved – Keep the Humanity in AI.