Most common failure modes for agents during QA
How we set up every coding agent to QA its own work: an isolated environment per change, realistic data, a written playbook, and a preflight.
How to maintain review quality with 10x PR volume
Four check categories, three outcomes per finding, and why a “required” AI review can report a pass when the model behind it is down.
Install Mesmer for Cursor: one engineer, a fleet, or a Build
Mesmer's Cursor integration installs with one shell command: once per machine for an individual engineer, as a recurring Rippling script policy plus a configuration profile for a fleet, or inside a Cloud Agent Build via environment.json. Capture every repository, no per-project opt-in, no per-engineer opt-out.
Most common agent errors
Five common ways teams produce wrong or stale context for their agents, and four principles that keep it from happening.
Agent analytics for Claude Code
A new AI Usage tab in Metrics: what your coding agents cost, who is using them, and how that spend tracks against delivered effort.
Connectors: bring your own MCP servers
Connect any OAuth-capable MCP server to Mesmer — pick from a vetted catalogue or paste a URL — and its tools work in chat and the workflow planner.
You may be spending more tokens on context gathering than writing code
Five context-gathering patterns that burn tokens in agentic eng orgs, and the fix for each: scoped instruction files, context firewalls, zero-token polling.
Golden standard for agent readiness: treat agents as new hires
We build harnesses for engineering orgs, and this is the metric we've seen is reliable at predicting increases in speed and quality developing with agents, proven through 8 months of experimentation including +2000 engineers across +100 companies. Measure the onboarding of an agent as if it were a new hire, looking at time from access to a deployed first feature, with no help.
New information architecture
A new Home page — your daily standup — and navigation that drops the org chart: pick a scope and timeframe once, and the other tabs all answer for it.
The playbook behind teams that 2x'd productivity in Q1
Hundreds of teams have connected their code and project data to Mesmer, and the ones that moved 2x+ faster in Q1 share a playbook. Humans stopped writing code directly, agents get triggered by alerts rather than by people, review moved off the human critical path in four stages, codebases were reshaped for retrieval, and pods shrank to two or three engineers per bet.
What I learned about hiring (and firing) scaling to 150
Hiring does not get easier as you scale, and firing fast is a fantasy for most founders — so build a company that hires slow instead. The mechanics that worked at 150 people: network downloads instead of "who do you know", mining interviews for names, one owner per hiring decision, 3 to 4 months of severance as policy, and quarterly skip levels.
Life after DORA: eng metrics for agentic teams
DORA still works as a health check but it is a lagging indicator you cannot optimize toward. SPACE is worse off — its volume components are game-able and it pins system problems on individuals. What replaces them: org-level measurement, volume used as a cross-reference rather than a target, quality graded on every merge, review coverage weighted by effort, and two per-person metrics.
How you compare across companies
Mesmer now plots each engineer's shipped effort per day against the Mesmer Benchmark — anonymized percentiles pooled from the hundreds of teams using Mesmer.
Compare: any two teams, people, or windows
Compare mode sets two scopes side by side — any team, person, or window against another — with a Normalized toggle so different-sized cohorts compare fairly.
Metrics page: the Quality tab
A new Quality tab on the Metrics page shows how your team reviews: effort reviewed, time to first review, rubber-stamp rate, and average PR quality.
Workflows: automate how you run your team
Workflows automate how you run your team — give feedback, catch stale work, flag metrics moving the wrong way, and build custom reports on your schedule.
Metrics page: zoom in
The Metrics page is the home for every engineering metric in Mesmer — Overview, Activity, Productivity, Delivery tabs, scoped to any team, any window.
Agentic PR attribution: crediting the human owner
When a coding agent leaves no co-author trace, Mesmer now credits the responsible human through a fallback chain: opener, then assignee, merger, approver.
Configurable weekends for teams with non-standard weeks
Choose which days count as non-working in Organization Settings, so Mesmer's metrics and AI assistant reflect your team's real work week.
Mesmer's MCP is ready for your agents
A Model Context Protocol server brings everything in Mesmer into the editors and assistants you already use — no dashboard required.
Briefings now have a project view
Mesmer's briefings now roll up by project — what shipped, what's in flight, next actions — with click-through to the contributions feeding each summary.
PR Quality grades on every merged change
Mesmer now grades every merged PR from 0 to 100, built from severity-tagged findings, normalized for PR size, and locked once calculated.
Know if your team's effort is seeing the light of day
Mesmer separates work that actually merged from PRs closed without merging, so effort and productivity reflect only what shipped.
Effort recalibration
Effort scores now use a wider range so easy work reads as easy and hard work can score higher than the old ceiling allowed.
