Introduction
Kraken is one of the world's leading multi-asset trading platforms globally, trusted by millions of institutions, professional traders, and consumers. The company moves fast by design, with thousands of engineers contributing across services, languages, and time zones.
"Everyone at Kraken is using AI right now, whether we like it or not," says Nik Sudan, who leads engineering operations at Kraken. AI has made it easier for engineers and non-engineers alike to spin up pilots and proofs of concept. Designers are more hands-on with code than before, often working in repositories separate from production.
All while industry-wide, AI spend per developer is climbing toward the cost of the developers themselves. Adoption metrics could not answer the question Nik cared about most - whether AI was improving the delivery pipeline. He needed proof that throughput was rising without quality or stability slipping.
Kraken turned to LinearB, the engineering productivity platform, as a way to measure AI's impact on engineering delivery, test assumptions about where work slows down, and translate delivery data for leaders who don’t live in the metrics every day.
Adoption metrics couldn't prove AI ROI
"AI usage isn't a metric that's great for effectiveness," says Nik. "It's one signal among many. But by itself, it doesn't prove anything about the SDLC and the value of AI."
Nik sees two mistakes leaders make. They treat spend as proof of value, assuming a team with a big AI bill must be getting their money's worth. Or they track adoption, counting seats filled and dashboards green, without asking whether the work got better.
Kraken wanted a harder standard. They wanted to know whether AI improves the work itself.
Balancing throughput, quality, and stability
Instead of a single AI metric, Kraken balances a portfolio of signals in LinearB across three dimensions.
- Throughput. Merge request output, frequency, and the size and maturity of those changes.
- Quality. Rework rate, bug escape rate, and how deep reviews actually go.
- Stability. Whether shipped work holds up in production, measured through incidents and regressions.
On top of those signals, Nik's team tracks what he considers the most useful measure available, which is cost per contribution. Divide AI spend by impact-weighted output, per merge, per engineer, or per sprint.

This data is all visible in LinearB, so Kraken can track cost per contribution without hopping between tools. "As people use AI more, their output should increase," Nik says.
Nik Sudan
Weighting by impact matters. Two engineers can produce the same raw output while one burns thousands of dollars in tokens and the other spends hundreds. Measured raw, the metric rewards whoever ships the most trivial changes. Weighted by impact, it rewards engineers who use the right model for the work, manage context effectively, and know when a problem calls for deterministic thinking rather than an agent.
The same discipline applies to how Kraken ships. Proofs of concept typically run in isolated, throwaway environments so the team can iterate quickly without dragging experiments straight into production. Designers prototype in separate repositories, removed from production mobile and web apps. When a concept is ready to ship, they hand working code to engineers, which Nik says is often easier to productionize than a Figma file or PDF.
When work crosses into the real delivery pipeline, LinearB keeps the rework rate visible so the team knows the speed isn't costing them stability.
Pulling the thread with LinearB to find the hidden bottleneck
Review time was one of Kraken's bottlenecks. Nik calls it the most important part of cycle time, and the hardest to explain. With thousands of engineers contributing daily across different services and languages, the cause could be anything.
The team had a hunch shared by most of the industry. Big, complex merge requests slow reviews down.
To test it, they used a two-layer approach. LinearB and its MCP server surface the high-level trends, slow teams, affected repositories, and relevant merge requests. Their Git provider provides granular details, such as reviewers, comments, files touched, and timestamps. AI moves fluidly between the two layers and joins the data, work that previously required a team of analysts and weeks of stitching.
As it turns out, the hunch was wrong. "Size and complexity accounted for maybe 5 to 10% of the slowdown," Nik says.
In just four months, Kraken’s engineering team worked hard to cut review time by over half, from 2 days, 15 hours to 1 day, 3 hours at the 90th percentile.

Making the data land with the business
Kraken pairs the analysis with deliberate translation for non-technical stakeholders.
The team prefers P90 over averages because averages can be misleading. A team doing well overall quietly absorbs the parts that are struggling, while P90 surfaces the slow tail, exactly what an engineering org with a high bar wants to see. Nik's team isolates each development group in LinearB and measures it against its own benchmark, never against a leaderboard.
When a team shows red against its benchmark, Nik looks for context before escalations reach leadership "A team at the bottom is usually one suffering from friction, or they need engineering investment," he says.
LinearB makes that context easy to share with non-technical stakeholders. Kraken uses the same delivery signals across multiple surfaces, depending on the audience:
- Executive reporting views for leadership-ready summaries
- The MCP server and API for deeper analysis and custom workflows
- The AI Insights dashboard for answering ad hoc questions
LinearB provides a shared language that leadership understands and helps Kraken engineering make the case for investments into tech debt and engineering-driven initiatives.
Use LinearB to find friction and bottlenecks early in your AI journey
Nik's advice for teams scaling AI comes down to one word.
Nik Sudan
For Kraken, that discipline turned AI measurement from a hunch into a standard their leaders trust. LinearB gives Nik's team the signal to see where delivery slows, the benchmarks to hold each team to its own bar, and the shared, trusted view that moves leadership from anecdote to action. As Kraken scales AI across thousands of engineers, LinearB helps answer the question they started with: Does AI improve Kraken's work with throughput, quality, and stability rising together?
Want the full story? Nik Sudan recently joined us on Dev Interrupted to go deeper on AI effectiveness, LinearB's MCP server, and the review bottleneck discovery that started as a hunch.