Devin Thomas strips AI guardrails for red teams
- Obliteration AI removes safety guardrails from open-source models to let security teams test vulnerabilities.
- A secondary software layer enforces corporate safety rules while keeping model parameters completely uncensored.
- Open-source Chinese models now trail US frontier models by roughly three months.
Devin Thomas is stripping the safety brakes off open-source artificial intelligence.
On This Week in Startups in October 2026, the Obliteration AI founder explained that commercial AI guardrails actively cripple enterprise cybersecurity defenses. When corporate red teams attempt to simulate cyberattacks or probe internal software agents for vulnerabilities, proprietary models from OpenAI and Anthropic issue blanket refusals. By using model interpretability to locate and alter the exact parameters responsible for refusals in open-source models like DeepSeek and GLM, Thomas removes those hardcoded barriers without degrading baseline reasoning performance.
The architectural breakthrough relies on an inversion of control. Instead of embedding rigid moral filtering deep within model weights, Obliteration AI strips the weights clean and routes user queries through an external, low-latency software policy layer. This secondary safety checker runs in milliseconds, enforcing custom enterprise security guidelines while preserving hard stops against self-harm or explicitly illegal acts.
Demand for uncensored weights has moved rapidly beyond niche defense contractors. Fortune 500 chief information security officers are now buying these modified models to attack their own autonomous agent deployments before external threat actors find the gaps. Anthropic recently underscored this risk shift by documenting the cyber attack efficacy of an obliterated GLM model, highlighting how quickly open-source weights can be weaponized or repurposed.
The competitive timeline between global AI developers is compressing fast. Thomas noted on the show that the performance gap separating open-source Chinese models from US frontier models has narrowed to approximately three months. As open-weight models catch up to proprietary systems, defensive security teams require equal access to uncensored capabilities to keep pace with potential adversaries.
Startups operating in this space face immediate structural bottlenecks. A persistent global GPU shortage restricts high-availability scaling for new security platforms, even as corporate demand accelerates. Data from venture firm Andreessen Horowitz shows that a massive percentage of enterprises have still not integrated AI, held back by lengthy internal approval chains and unresolved security risks.
Beyond technical guardrails, host Jason Calacanis addressed broader predictions about AI replacing professional service business models. While critics argue automated agents will eliminate hourly billing in legal and accounting fields, Calacanis insisted clients pay human experts primarily for liability and risk transfer during litigation. Automated tools merely push advisers upstream to deliver higher-value work per hour.
A similar dynamic protects traditional startup founding teams. Despite tools enabling solo builders to write code faster, venture capitalists continue to back multi-founder teams to avoid single points of failure. Software handles execution, but humans remain necessary to share organizational risk and manage operational pressure.
Source Intelligence
- Deep dive into what was said in the episodes
Inside The Startup Building Uncensored AI (Abliteration AI) | EP 2345 • Oct 3
- Devin Thomas explains that Obliteration AI removes safety guardrails from open-source models to serve industries like cybersecurity and defense that are blocked by frontier labs. The platform allows companies to run unrestricted models natively.
- Devin Thomas introduces an inversion of control feature that separates the unrestricted LLM from a customizable software policy layer. This architecture allows organizations to implement their own security guidelines without sacrificing model performance or adding significant latency.
- Devin Thomas explains that the obliteration process uses model interpretability to locate and modify the specific parameters responsible for model refusals. A secondary, safety-focused model then runs in milliseconds to handle user-defined compliance policies.
- Devin Thomas claims the performance gap between open-source Chinese models and US frontier models has closed to approximately three months. Anthropic highlighted this capability shift by documenting the cyber attack efficacy of an obliterated GLM model.
- Devin Thomas reports that Fortune 500 Chief Information Security Officers are adopting uncensored models to proactively red-team their own systems. However, Thomas notes that a real GPU shortage limits high-availability scaling for startups in this space.
- Devin Thomas cites an Andreessen Horowitz report indicating that a massive percentage of enterprises have not yet adopted artificial intelligence. Thomas suggests complex corporate approval processes and security concerns are stalling widespread enterprise integration.
Also discussed on this episode: (4)
Labor (1)
- Jason Calacanis argues AI will not eliminate the billable hour in professional services, mirroring how previous technological shifts failed to displace it. Instead, Calacanis predicts professionals will move upstream, using AI to deliver better work faster while charging for liability and trust.
Media (2)
- Bruno Villela describes the rapid growth of microdramas, which split two hours of content into dozens of short episodes. Villela notes creators use ComfyUI and C-dance 2.5 to build standardized pipelines that automate and scale video production.
- Jason Calacanis predicts the emergence of an AI-driven Toy Story moment. Calacanis expects AI video generation to merge with traditional production tools like ComfyUI, allowing creators to precise-control environments and pre-sketched characters rather than relying on basic text prompts.
VC (1)
- Jason Calacanis asserts that venture capitalists still heavily favor multi-founder teams over solo founders despite AI-driven productivity gains. Calacanis argues multiple founders are necessary to handle extreme task-switching and mitigate the existential risk of a single founder quitting.
