Diogo Almeida warns AI coding agents fail at app logic
- Vibe coding dropped simple app prototyping costs from $500,000 to $5,000 in two weeks.
- Engineers warn AI coding agents automate raw syntax generation without improving underlying software logic.
- Building reliable software requires structured decision primitives rather than non-deterministic AI external wrappers.
AI coding agents write syntax faster than ever, but software isn't getting smarter.
On This Week in Startups, host Jason Calacanis demonstrated how dramatically development costs have dropped for surface-level software. A browser extension that would have cost $500,000 two decades ago - and $50,000 just two years ago - was built in under two weeks through a $5,000 bounty. Developers like Robert shipped functional apps such as annotated.wtf using AI prompt-driven development, known as vibe coding.
Yet that rapid prototyping reveals immediate boundaries. Submissions like Chirag Azarpoda’s web annotation tool stumbled on clunky user interfaces and manual entry steps, demonstrating that generating raw code is fundamentally different from designing intuitive product workflows. Calacanis noted that while he plans to incubate winning bounty submissions into venture projects, speed alone does not build durable enterprise value.
That breakdown at the interface reflects a deeper architecture problem across the industry. On The a16z Show, TypeSafe founder Diogo Almeida argued that tools like Claude Code and Cursor merely accelerate standard syntax generation. They act as faster typists without modifying underlying application logic, leaving software programs structurally identical to those built a decade ago.
Co-host Martin Casado highlighted internal data from Google showing that the average pull request at major tech firms spans just ten lines of code. Accelerating raw typing targets a minor operational bottleneck rather than the hard engineering challenges of software reliability. Co-host Ben Horowitz further warned that dumping unmonitored, AI-generated code into existing codebases risks multiplying security vulnerabilities at scale.
Almeida emphasized that despite continuous progress on artificial benchmarks, major AI labs have tried and failed to automate basic customer support workflows since 2020. The disconnect stems from chasing artificial general intelligence milestones instead of providing tight statistical guarantees. Software engineers require predictable behavior, whereas non-deterministic AI outputs create chaotic edge cases when handling corporate workflows.
Rather than relying on external agent wrappers to write imperative code, Almeida advocates embedding language-driven decision nodes directly inside existing state machines. His company developed Jev, a structured classifier designed to work like a standard database library. It allows developers to trade off execution budget, speed, and intelligence while handling non-deterministic inputs cleanly.
Vibe coding makes simple apps cheap, but enterprise software demands operational dependability. Until AI shifts from drafting lines of syntax to executing reliable logic, the true cost of software automation remains unpaid.
Source Intelligence
- Deep dive into what was said in the episodes
Jason’s Put a $5K Bounty on His Dream Chrome Extension | E2338 • Sep 28
- Jason Calacanis argues vibe coding has collapsed development costs. A Chrome extension requiring $500,000 to build 20 years ago, and $50,000 two years ago, can now be prototyped for a $5,000 bounty within two weeks.
Also discussed on this episode: (10)
Startups (4)
- Jason Calacanis purchased the domain annotated.com for $30,000 to $40,000 after selling Weblogs, Inc. to AOL. His goal was to build a billion dollar startup allowing users to highlight, dispute, or discuss web text.
- Alan Shifflett's submission received an 8 out of 10 score. Jason Calacanis praised its categorized tags for annotations, such as steelmanning or fact checking, and its robust social profile sidebars, despite minor visual clunkiness.
- Chirag Azarpoda submitted a version using a misspelled domain to capture typo traffic. It received a 6 out of 10 score due to its manual, step heavy timestamp clipping process, which Jason Calacanis labeled as poor user experience.
- Jason Calacanis asserts that product demos succeed or fail based on the quality of their examples. Founders must avoid janky or simplified examples, using realistic, highly detailed use cases to make their product's value instantly clear.
Regulation (1)
- Robert's submission, annotated.wtf, scored 8.5 out of 10 for its clean video clipping tool. Lon and Jason Calacanis agree that sharing short, multi minute clips for criticism falls under fair use and has become culturally accepted.
VC (2)
- Jason Calacanis outlines how venture funds can incubate startups internally by allocating capital from large funds to build concepts from scratch. The general partners hire external management teams, sharing equity with them and limited partners.
- Jason Calacanis advises a homeless viewer with 700 pages of business plans to secure stable employment and housing before pursuing venture funding. He emphasizes that venture capitalists are highly unlikely to fund homeless founders.
Models (1)
- Jason Calacanis suggests that web annotation platforms can monetize user data by selling expert annotated content feeds to AI companies. These human curated datasets provide the essential domain context that standard LLM training datasets currently lack.
Psychology (1)
- Jason Calacanis cites the Pygmalion effect to emphasize that high societal expectations can positively influence an individual's performance and self belief, encouraging the homeless viewer to maintain their ambition while stabilizing their life.
Diplomacy (1)
- Jason Calacanis and Lon explain that although sending governments maintain embassies, high profile ambassadors must personally finance elaborate receptions and state dinners. These out of pocket diplomacy expenses often reach millions of dollars annually.
AI Can Write Code. Why Isn’t Software Better? • Sep 28
- Diogo Almeida argues coding agents like Cursor or Claude Code merely speed up the production of traditional code without making software smarter. Jev provides a new programmatic primitive that integrates natural language intent directly into program state decisions.
- Diogo Almeida describes Jev as a highly advanced classifier designed to function like a database or standard library within software systems. It accepts natural language inputs to guide state machine decisions, avoiding floats where AI predictably struggles.
- Martin Casado notes a Google study revealed the average pull request at a large tech company is only ten lines long. This suggests that accelerating syntax generation via AI agents targets a minor bottleneck in software development.
- Diogo Almeida anticipates a new era of probabilistic programming where developers trade off intelligence, cost, and speed. Systems engineers will use lightweight, approximate guesses to dynamically route data within complex infrastructure.
Also discussed on this episode: (6)
Startups (1)
- TypeSafe prioritizes intelligence per dollar over speed as its primary design metric. Diogo Almeida explains this focus is necessary to ensure AI can be integrated deeply into the internal state of critical software systems rather than just human-facing chat layers.
Models (3)
- Diogo Almeida entered the AI field by winning a Kaggle competition using system automation and nested loops rather than advanced mathematics. This victory led SVM co-inventor Isabel Guyon to introduce him to the broader machine learning research community.
- At OpenAI in late 2021, Diogo Almeida realized RLHF generalization was real after testing the model with the query "why is it important to eat socks before meditating." The model generated plausible, human-like answers for the unseen query.
- Diogo Almeida rejects the idea that AI is on a path toward recursive self-improvement. However, he believes automating economically valuable work is highly achievable because the majority of real-world work consists of simple, rote instructions.
Agents (1)
- OpenAI has attempted to automate customer service since 2020 without achieving robust reliability. Diogo Almeida attributes this failure to the industry optimizing models for human judges and impressive demos rather than background task execution.
Enterprise (1)
- The Saas-pocalypse theory failed because complex software architecture is extremely difficult to replicate beneath the surface. Diogo Almeida predicts an inverse Saas-pocalypse where established SaaS companies become the biggest AI winners by deeply automating existing customer workflows.

