What Top 1% of Agentic Engineers Do Differently
Summary
This video discusses the rapid advancement and adoption of AI agents, focusing on token consumption, efficiency, and the future of AI development. It highlights the massive scale of token usage, the surprising number of 'token billionaires,' and the ongoing debate between large and small models. Key themes include embracing 'slop' for innovation, the evolution of developer workflows with agents, the importance of mission-driven leadership, and the strategic considerations for AI startups, from domain-specific agent labs to foundational model research. The discussion also touches on the need for human taste and judgment in product development and the future of AI infrastructure and capabilities, emphasizing faster inference as a major upcoming advancement.
Key Insights
Embracing 'slop' and experimentation is crucial for discovering innovative uses of AI agents.
While efficiency is important, engineers who are willing to explore and experiment with 'slop' or less refined outputs from AI agents are currently underrated. This non-consensus approach is key to finding groundbreaking applications.
The frontier of AI will always see large models, but routine tasks create space for specialized models and agent labs.
While major labs focus on large, frontier models (like GPT-5.6), most AI work involves routine tasks. This creates opportunities for agent labs like Cognition, which focus on domain-specific problems (e.g., enterprise coding) and may train their own models, and for smaller, specialized models.
The value of large models is unlocked by exploring novel prompts and abstract concepts, not just by incremental improvements.
Simply using old prompts with new, larger models yields only marginal improvements. True benefit comes from experimenting with diverse prompts and asking open-ended questions to explore the model's unknown capabilities, differentiating expert users from basic tool users.
User testing and real human feedback are indispensable, even with massive AI token usage, to avoid building unusable products.
Despite high token consumption, products must be validated by real users. It's easy to generate 'slop' that makes an app unusable due to AI's incentivization towards token usage. Without genuine user needs driving development, startups risk financial failure.
Sections
AI Agent Adoption and Token Usage
AI engineers are extensively using AI, with 'token billionaires' program attracting far more participants than expected.
The 'token billionaire' program, rewarding those who spend a billion tokens a week, saw around 300 participants out of 7,000 attendees, indicating a high level of AI adoption and experimentation. Some companies like Cursor are reportedly using 8-10 billion tokens per day.
Embracing 'slop' and experimentation is crucial for discovering innovative uses of AI agents.
While efficiency is important, engineers who are willing to explore and experiment with 'slop' or less refined outputs from AI agents are currently underrated. This non-consensus approach is key to finding groundbreaking applications.
Career opportunities exist for early adopters exploring innovative agent use cases, despite potential initial inefficiency.
Investing time in exploring AI agents is seen as a good career bet because there's no downside, companies cover the costs, and early success in finding innovative uses can yield significant career upside. Many initial applications might be inefficient but exploring them is valuable.
Workflow Changes and Tooling
Developers are shifting focus from managing multiple agents to maintaining concentration on single, high-impact tasks.
A personal workflow change observed is consciously focusing on one primary task rather than running numerous agents in parallel on different tasks. However, other background tasks for repetitive actions, research, or prototyping are still valuable.
The need for traditional IDEs is declining as agents handle more complex coding and review tasks.
Tools like Cognition's Devon are integrating AI deeply into workflows, reducing the reliance on traditional IDEs. The IDE's role is shifting towards being primarily a file editor or a tool for reviewing code, with the potential for cloud-based or agent-driven environments to take over.
Graphical user interfaces for agents can be helpful but depend on user preference; many prefer simpler list or thread-based visualizations.
While graphical interfaces like Warcraft-style simulations can aid visualization for some users, most people are comfortable with independent threads or flat lists. A middle ground between flat lists and full graphical interfaces, like a conductor or desktop approach, seems practical.
Model Development and Future Trends
The frontier of AI will always see large models, but routine tasks create space for specialized models and agent labs.
While major labs focus on large, frontier models (like GPT-5.6), most AI work involves routine tasks. This creates opportunities for agent labs like Cognition, which focus on domain-specific problems (e.g., enterprise coding) and may train their own models, and for smaller, specialized models.
Post-training compute is becoming as significant as pre-training compute for advanced AI models.
The compute required for post-training, such as fine-tuning and continued training, is substantial, sometimes equivalent to pre-training. This blurs the lines between fine-tuning and continued training, requiring significant resources.
Large models offer significant advantages, and claims of small models' superiority often stem from practical limitations or sales pitches.
While some advocate for small models, proponents of large models argue that many such claims are either attempts to sell startups or a result of lacking the necessary GPU resources to run large models. The trend suggests continued development and use of very large parameter models (10-20 trillion parameters).
The value of large models is unlocked by exploring novel prompts and abstract concepts, not just by incremental improvements.
Simply using old prompts with new, larger models yields only marginal improvements. True benefit comes from experimenting with diverse prompts and asking open-ended questions to explore the model's unknown capabilities, differentiating expert users from basic tool users.
Agent Loops and Development best practices
Effective agent loops require clear specifications, verification criteria, and explicit 'what not to do' instructions.
A great agent loop needs to define desired outcomes, provide verification for completion, and importantly, specify actions to avoid (e.g., generating overly large files, neglecting mobile responsiveness) to guide the agent effectively and prevent common pitfalls.
Developing a deep understanding of data structures is crucial for managing complex AI projects and avoiding wasted effort.
When working on significant projects, understanding the data layer—how data is logged, stored, accessed, and represented—is paramount. This understanding prevents issues with codebases becoming unmanageable and allows for effective rollback or refactoring.
User testing and real human feedback are indispensable, even with massive AI token usage, to avoid building unusable products.
Despite high token consumption, products must be validated by real users. It's easy to generate 'slop' that makes an app unusable due to AI's incentivization towards token usage. Without genuine user needs driving development, startups risk financial failure.
Founders with 'taste' prioritize the problem and industry progress over self-promotion and are transparent about their claims.
Founders with taste are genuine, avoid overstating capabilities (e.g., small models beating large ones), focus on contributing to the industry's collective progress, and share insights strategically rather than just marketing their product. They connect their work to a broader mission.
Developing a strong, mission-driven narrative is essential for rallying support and attracting talent and customers.
A compelling mission, focused on a broader societal or industry goal (like advancing human civilization or solving a major problem), inspires support. In contrast, a narrow focus on beating competitors or simply accumulating wealth doesn't resonate as strongly.
Starting an AI Startup and Future Outlook
AI startups can succeed by focusing on domain expertise within specific verticals or by pursuing foundational model research.
Two primary paths exist: building 'agent labs' that master a vertical (coding, healthcare) and remain model-agnostic, or engaging in high-risk, high-reward domain-specific model research. Agent labs are less risky and pivotable, while model labs attract more capital but face higher failure rates.
Ambitious founders should build substantial infrastructure and deploy capital strategically, not shy away from large investments.
Building true infrastructure (data centers, GPUs, data) and deploying significant capital for a competitive advantage is key for ambitious founders. Those who are timid, focusing only on open-source frameworks or consulting, are unlikely to achieve large-scale success.
Thinking on longer timelines (10-30 years) naturally fosters greater ambition than focusing on short-term (2-5 years) goals.
Increasing the time horizon for planning and ambition is a strategy to encourage larger-scale thinking. Building for what will be needed decades from now, rather than just the next few years, helps overcome timidity and identify significant opportunities.
The development of custom silicon and advancements in agent infrastructure are key areas for ambitious AI innovation.
Companies are exploring custom silicon to optimize for AI workloads beyond general-purpose GPUs, and reinventing serverless/cloud paradigms for agents (e.g., Modal, E2B, Daytona) to improve computer spin-up times and deployment efficiency.
Almost everything built before the AI explosion has an opportunity for disruption or optimization through AI integration.
Existing technologies and systems, from software protocols to user interfaces, can be re-evaluated and reinvented with AI capabilities. While fundamental internet layers might remain sticky, surface-level applications and communication protocols are ripe for transformation.
Building for both human users and agents is important, as human-centric design often benefits AI interactions.
The distinction between building for humans versus agents is often false. Good human design principles can benefit agent interactions, and vice versa. Starting with human-centric features may be practical for new products since AI integration takes time to scale.
The definition of 'individual tasks' is elevating to more abstract levels, driven by powerful agents like Codex.
Interacting with agents often means delegating tasks at a higher level of abstraction (e.g., 'pay my medical bill' with just the website URL). This shift promises to automate routine white-collar work, fostering creativity and lowering activation energy for exploration.
Specialization is key to being at the cutting edge of AI, but a broad service layer for specialists is also vital.
Most individuals should aim to be specialists in a niche (e.g., memory, generative media). A smaller group, perhaps with broader interests, can provide the connecting 'service layer' that enables specialists to go deep.
High-performance teams benefit from a mix of experienced supervisors and creative, less experienced individuals.
The ideal team composition includes experienced supervisors to guide and less experienced but 'spiky' individuals for creativity. Practical hiring involves work trials, regular feedback, and prompt, yet fair, management of underperformers to maintain team standards.
Super-fast inference (thousands of tokens per second) is imminent and will fundamentally change product possibilities.
Current inference speeds are rapidly increasing, with demos showing thousands or even hundreds of thousands of tokens per second. Productionizing these speeds will enable entirely new types of AI-powered products and applications.
Ask a Question
*Uses 1 Wisdom coin from your coin balance









