Clients

The work

Most of this work happened under NDA, inside regulated environments, or close to a competitive roadmap - so the case studies are anonymized. Each one shows what the situation was, what we were brought in to do, what we actually did, and what was different when we left.

Case Study 01

The first AI workload at a heavily regulated company

Context. A heavily regulated company with no AI in production. The company was writing its first AI governance framework at the same time its first AI workload was being built - the rules and the system that had to pass them were developed in parallel, which is the situation most regulated enterprises actually face.

Engagement. Engineering leadership for the team building that first workload - and ownership of getting it through every review standing between a working system and production.

What we did. Shipped the company's first production AI/ML workload, on AWS Bedrock. The model integration was the straightforward part. The work was clearing the workload through risk, legal, security, and audit review, and working through PII and financial data handling requirements - in an environment where each of those functions was evaluating AI for the first time, against a governance framework a central AI committee was still drafting. That meant making the system legible to non-engineers: what data goes in, where it flows, what the model can and cannot do with it, and how that is enforced in the architecture rather than promised in a document.

What changed. The workload cleared review and went to production. More durable than the workload itself: the company now had a worked example - a first AI system that had survived risk, legal, security, and audit - setting the precedent and the path for every workload after it. This engagement is why we tell clients that governance, not models, is the hard part of enterprise AI.

Case Study 02

Honing AI-augmented engineering

Context. A small B2B SaaS engineering team building at full AI-augmented velocity. The question was not whether AI made them faster - it visibly did. The question was whether anyone could see what that speed was doing to the codebase, and when the bill would come due.

Engagement. We advised the team on measurement and workflow: instrument the work, read the data, and tune how much autonomy the AI gets - deliberately, per phase of work, instead of by default.

What we did. Measured the work at the PR level: roughly 48,000 absolute lines over 15 days, with throughput holding constant at about 3,000–4,000 absolute lines per day across every phase. The add/remove ratio turned out to be the telemetry. Feature bursts ran high ratios; consolidation ran low - the same engine in a different gear, visible in data the team already had in git. From that data, stabilization stopped being a feeling and became a schedule: after roughly a week of high-ratio feature velocity, the team ran a dedicated cleanup pass, reverse-engineered the recurring mistakes into written standards, and fed those standards back into the AI's persistent context so the next cycle started cleaner than the last.

Cleanup pass234551015Dayk lines per day · add:remove ratio
Absolute lines per day (thousands)
Add:remove ratio

What changed. The team now reads burn versus replenish directly off its git history, and AI autonomy is a dial they set per work phase rather than a setting they forgot they chose. The frameworks on this site - AI Code Nines, the Quality Drawdown, the Stabilization Pass - were extracted from this engagement with a client engineering team we advised.

Case Study 03

Building a new AI product line

Context. An enterprise technology company with an established platform, a decision to build a new AI-driven product line on top of it, and a parallel track of acquiring AI-native companies to accelerate the move. Strategy, architecture, and M&A were happening at the same time, and each one constrained the other two.

Engagement. AI strategy and architecture, hands-on production build, and technical due diligence on acquisition targets - the same consultant across all three, from product definition through production code. The path from product idea to working technology was the engagement; there was no handoff between the person setting direction and the person writing the reference implementation.

What we did. Defined the AI/ML platform roadmap and how the new capability integrates with the existing platform. Built the new product's AI architecture hands-on: embeddings on pgvector, clustering and topic extraction, and RAG that mixes vector retrieval with structured relational data - because the right context for a model is rarely just the nearest neighbors. Established reference patterns for their engineering team, so the architecture survives the consultant leaving. In parallel, ran M&A technical due diligence on AI-native targets: architecture review, platform and team capability assessment, integration risk.

What changed. The AI layer is in their production codebase, and their engineers extend it without us.

How we engage

Conversation first

The first conversation is a working session: what you're building, where AI fits or doesn't, and what is actually blocking you.

Engagements take the shape of the problem: hands-on architecture and build, fractional leadership, diligence on a deal, or enablement for a team that's already moving. Everything we recommend is something we have shipped.