Skip to content

Tag

#cto

  • Most common failure modes for agents during QA

    How we set up every coding agent to QA its own work: an isolated environment per change, realistic data, a written playbook, and a preflight.

    Sergio Bergmann

  • How to maintain review quality with 10x PR volume

    Four check categories, three outcomes per finding, and why a “required” AI review can report a pass when the model behind it is down.

    Sergio Bergmann

  • Most common agent errors

    Five common ways teams produce wrong or stale context for their agents, and four principles that keep it from happening.

    Sergio Bergmann

  • You may be spending more tokens on context gathering than writing code

    Five context-gathering patterns that burn tokens in agentic eng orgs, and the fix for each: scoped instruction files, context firewalls, zero-token polling.

    Sergio Bergmann

  • Golden standard for agent readiness: treat agents as new hires

    We build harnesses for engineering orgs, and this is the metric we've seen is reliable at predicting increases in speed and quality developing with agents, proven through 8 months of experimentation including +2000 engineers across +100 companies. Measure the onboarding of an agent as if it were a new hire, looking at time from access to a deployed first feature, with no help.

    Sergio Bergmann

  • The playbook behind teams that 2x'd productivity in Q1

    Hundreds of teams have connected their code and project data to Mesmer, and the ones that moved 2x+ faster in Q1 share a playbook. Humans stopped writing code directly, agents get triggered by alerts rather than by people, review moved off the human critical path in four stages, codebases were reshaped for retrieval, and pods shrank to two or three engineers per bet.

    João de Paula

  • What I learned about hiring (and firing) scaling to 150

    Hiring does not get easier as you scale, and firing fast is a fantasy for most founders — so build a company that hires slow instead. The mechanics that worked at 150 people: network downloads instead of "who do you know", mining interviews for names, one owner per hiring decision, 3 to 4 months of severance as policy, and quarterly skip levels.

    João de Paula

  • Life after DORA: eng metrics for agentic teams

    DORA still works as a health check but it is a lagging indicator you cannot optimize toward. SPACE is worse off — its volume components are game-able and it pins system problems on individuals. What replaces them: org-level measurement, volume used as a cross-reference rather than a target, quality graded on every merge, review coverage weighted by effort, and two per-person metrics.

    Lucas Silva