From the lab
Field notes
Practical perspectives on AI engineering, deployment, and strategy from our work with enterprise clients.
How to Evaluate AI Coding Agents Before They Touch Production
Benchmark scores like SWE-bench tell you less than you think. Here is the evaluation stack we recommend — public benchmarks, private evals on your own codebase, and hard rollback guardrails — before an AI coding agent's output ships to production.
Generative Engine Optimisation (GEO): How to Win AI Search in 2026
AI answer engines like ChatGPT, Perplexity and Google AI Overviews now decide which companies get recommended — and they cite only a handful of sources per answer. Here is what replaces classic SEO tactics, backed by the research, and a practical 90-day plan to get your firm cited.
Human-in-the-Loop at 10x Speed: AI Code Review Best Practices
AI has made code generation nearly free, which makes review the real constraint. Here is the 2026 playbook for keeping a named senior engineer accountable for every merge: risk-tiered review, capped pull requests, machine first passes, and instrumented quality gates.
Where AI Cuts Software Delivery Costs by 50% — and Where It Doesn't
An honest, data-backed breakdown of AI's cost impact across the SDLC — specs, build, test, review and ops — using findings from DORA 2025, Stanford's 2026 AI Index, METR's randomised trial and GitClear's code-quality research.
MCP in Production: Security, Auth and Tool Governance Lessons
Hard-won lessons from integrating Model Context Protocol servers into enterprise systems: OAuth 2.1 done properly, defences against tool poisoning, and the governance layer most teams forget to build.
Spec-Driven Development: Why the Spec Is Becoming the Source Code
AI agents can now write most of the code — which makes the specification the artefact worth engineering. How Spec Kit, Kiro and the spec-driven workflow turn requirements into working, verifiable software.
The AI-Native Software Agency: What Changes When Agents Write the Code
When coding agents write most of the code, the software agency changes shape: small senior pods replace delivery pyramids, fixed outcomes replace billable hours, and verification becomes the product. Here is what actually changes — and what clients should demand in 2026.
When AI Makes Sense: A Framework for Enterprise Decision-Making
Not every problem needs AI. We share our evaluation framework for determining when machine learning adds genuine value versus when simpler solutions suffice.
Building Reliable AI Systems for Defence Applications
Mission-critical environments demand a different approach to AI deployment. Lessons learned from building systems where failure is not an option.
The Hidden Costs of AI Projects: What Nobody Tells You
Beyond compute and talent, AI initiatives carry costs that rarely appear in business cases. A realistic look at total cost of ownership.
Computer Vision in Healthcare: Augmenting, Not Replacing, Expertise
How we designed diagnostic assistance tools that enhance clinician decision-making while maintaining full accountability and transparency.
From POC to Production: Closing the AI Deployment Gap
90% of AI projects never make it to production. We examine the technical and organisational factors that determine success.
Data Quality Over Data Quantity: Building Effective Training Pipelines
In our experience, curated datasets consistently outperform larger, noisier alternatives. A practical guide to data-centric AI development.