Custom AI Software Built Around Your Business.
When off-the-shelf tools cannot express how your business actually works, the sensible answer is to make the software fit the business rather than reshape the business to fit the software. We design, engineer and hand over AI products that run on your infrastructure and follow your rules.
- Product-grade engineering
- Model-agnostic architecture
- You own everything
Eleven product shapes we have engineering patterns for
These are not service categories invented for a website. Each one has a known architecture, a known set of failure modes and a known way to test it, which is why they can be quoted and delivered instead of explored.
AI SaaS products
Multi-tenant applications with account isolation, usage metering, role-based access and per-plan model limits — built to be sold to customers, not demoed to a board.
AI dashboards
Operational views where each chart arrives with a written explanation of what moved, generated from the same query that drew the chart.
AI portals
Client, supplier or patient portals where the other side uploads documents, asks questions and gets grounded answers instead of emailing your team.
AI-powered web applications
Full products where the AI is a working part of the flow: drafting, classifying, extracting, routing and deciding inside the screens people already use.
AI mobile applications
iOS and Android apps with offline-tolerant queues, background sync, push notifications and on-device capture feeding the same AI services as the web app.
Internal AI platforms
A shared internal layer — auth, data access, retrieval, prompt registry, evaluation and logging — so every future AI feature ships in days instead of restarting from zero.
AI analytics
Natural-language querying over your warehouse with generated SQL shown to the user, plus scheduled narrative summaries of what changed since last period.
Recommendation systems
Ranking for products, content, next actions or candidates, combining behavioural signals with business rules so commercial constraints are never overridden by a model.
AI search
Semantic and hybrid search across your documents, tickets and records, with permission filtering applied before retrieval rather than after generation.
Document intelligence
Extraction from invoices, contracts, forms and scans with per-field confidence scores, a human review queue for anything below threshold, and a full audit trail.
Custom AI assistants
Assistants embedded in your product or back office that can call your APIs, respect your approval rules and hand off to a person with the full conversation attached.
Idea → Architecture → Development → Testing → Deployment → Optimisation
Six phases, each ending in something you can hold: a document, a running environment, a score, a release. Durations below are relative rather than promised — real timelines come from the scope agreed in phase one.
Build lifecycle
live workflow- IdeaScope + criteria
- ArchitectureDecide before code
- DevelopmentTwo-week cycles
- TestingTests + evaluations
- DeploymentStaged rollout
- OptimisationQuality + cost
Phases overlap in practice. Testing starts in the first development cycle, and optimisation begins the day real traffic arrives.
Idea
We turn a description of the problem into a scoped product definition: who uses it, what decision it changes, what it must never do, and how we will know it works. Anything that cannot be tied to a user or a cost is cut here rather than discovered in month three.
You receive
- Scoped product definition
- Success criteria
- Explicit non-goals
Architecture
Data model, service boundaries, integration points, retrieval strategy, model routing and guardrails are decided before code. We also decide what stays deterministic — the parts of the flow that should be plain code, not a model.
You receive
- System architecture
- Data + retrieval design
- Integration and guardrail plan
Development
Engineering runs in short cycles, each ending in something you can open and use on your own data. Length depends on surface area rather than a fixed number: a single workflow tool is short, a multi-tenant platform with billing and mobile is not.
You receive
- Working software each cycle
- Environment you can log into
- Cycle demo and changelog
Testing
Two layers. Conventional tests for the application, and an evaluation suite for the AI behaviour: a fixed set of real inputs with expected outcomes, scored on every change so a prompt or model swap cannot quietly degrade quality.
You receive
- Test suite
- Evaluation set with scores
- Known-limitations list
Deployment
Infrastructure as code, staged rollout, monitoring and alerts wired before the first real user, and a documented rollback. Where the AI touches customers, we usually start with a human approving output and remove that gate once the numbers earn it.
You receive
- Production environment
- Monitoring + alerting
- Rollback and runbook
Optimisation
Live systems drift as inputs, users and models change. We track quality, latency and cost per outcome, then tune prompts, retrieval, routing and model selection against what real usage shows rather than what the demo suggested.
You receive
- Quality + cost reporting
- Tuning cycles
- Roadmap for the next release
The parts of AI software that decide whether it survives contact with users
Most AI projects do not fail at the demo. They fail six weeks later, when quality quietly slips, cost climbs, nobody can explain a wrong answer and the only person who understands the system has moved on. These six standards exist specifically to prevent that.
Architecture that stays portable
Model calls sit behind an internal interface, prompts live in version control, and business logic never depends on one provider’s response shape. The application is a normal, well-structured system that happens to call AI services — so it can be maintained by any competent engineering team, including yours.
- Modular services
- Provider adapters
- Infrastructure as code
Evaluation, because unit tests miss regressions
A passing test suite tells you the code runs; it says nothing about whether answers got worse. We build a labelled evaluation set from your real inputs and score every change on accuracy, grounding, refusal behaviour and format compliance, with a threshold that blocks release when quality drops.
- Golden datasets
- Regression gates in CI
- Human review sampling
Observability at the trace level
Every AI interaction is logged as a trace: inputs, retrieved context, model and version, tokens, latency, cost and outcome. When a user says the system was wrong on Tuesday, we can open that exact request and see why, instead of guessing from an aggregate dashboard.
- Request-level tracing
- Quality + latency dashboards
- Alerting on drift
Security and access control
Permissions are enforced at retrieval, so a model can only reason over documents the current user is already entitled to read. Secrets are managed, PII handling is explicit, prompt-injection defences are applied at tool boundaries, and every privileged action is written to an audit log.
- Permission-aware retrieval
- Least-privilege tool access
- Auditable actions
Cost control by design
Inference cost is a product decision, not a surprise invoice. We budget tokens per request, route easy work to smaller models and reserve large ones for hard cases, cache aggressively where inputs repeat, and set per-tenant limits so one heavy user cannot consume the account.
- Model routing by difficulty
- Caching + batching
- Per-tenant budgets and alerts
Handover you can actually use
Delivery includes architecture documentation, runbooks for the failure modes we hit during build, environment setup that a new engineer can follow in an afternoon, and recorded walkthroughs. We measure this by whether your team can ship the next feature without calling us.
- Architecture + decision records
- Runbooks
- Engineer onboarding guide
Model-agnostic by default, because the best model keeps changing
The model you would pick today is unlikely to be the model you want in a year. Pricing shifts, context windows grow, a smaller model becomes good enough for two thirds of your traffic, and a regulator or a client asks where inference happens. Architecture is what decides whether any of that is a configuration change or a rewrite.
One interface, many providers
Application code calls an internal capability — summarise, extract, classify, rank — never a vendor SDK directly. Swapping the model behind that capability is a configuration change plus an evaluation run, not a refactor.
Prompts and schemas are versioned artefacts
Prompts, output schemas and retrieval settings live in version control with the code that uses them, so a change is reviewable, testable and revertible like any other change.
Routing by task, not by loyalty
Cheap, high-volume classification and routing go to small models. Reasoning-heavy work goes to larger ones. Sensitive workloads can be pinned to a self-hosted model without touching the product surface.
Evaluation makes the swap safe
Because the evaluation suite is built on your data, a model change is a measurable decision: run the suite, compare accuracy, latency and cost, then decide. That is the difference between switching models and hoping.
Typical building blocks
The specific technologies are chosen per project against your constraints — hosting region, existing stack, in-house skills and compliance. These are the roles that need filling in almost every AI product we build.
- REST and GraphQL APIs
- Vector stores
- Relational databases
- Object storage
- Job queues
- Schedulers
- Authentication and SSO
- Role-based access control
- Caching layers
- Webhooks and event streams
- Observability and tracing
- Feature flags
We are not a reseller and hold no vendor quota to fill. Where a commercial platform is the right answer we will recommend it, and where it is only the convenient answer we will say that too.
Already running AI somewhere?
Custom software rarely lands on an empty desk. If the goal is to connect a new build to the CRM, ERP and helpdesk you already run, that integration work is a discipline of its own.
See AI IntegrationsWhat you own at the end
A custom build is only worth more than a subscription if you genuinely end up owning it. Everything below transfers to you, and the handover is treated as a deliverable rather than a courtesy.
Source code
The full repository with its history, in your organisation, under a licence that gives you unrestricted commercial use.
Infrastructure
Cloud accounts, environments and deployment pipelines in your name, defined as code so they can be rebuilt without us.
Data
Your records, documents, embeddings and logs stay in storage you control, exportable at any time in an open format.
Prompts and evaluation suites
The prompt library and the labelled evaluation sets built from your data — the part that is genuinely hard to recreate.
Documentation
Architecture notes, decision records, runbooks and walkthroughs written for an engineer who has never met us.
What teams ask before commissioning a build
If the answer you need is not here, bring it to the call. We would rather scope it properly than guess in public.
Buy when the process is standard and a mature product already models it well — accounting, payroll, generic helpdesk. Build when the way you do the work is the advantage, when no product covers the specific sequence your business runs, when licence cost scales badly against the value you get, or when the data you would have to hand a vendor is the sensitive part. In practice most engagements are a mix: buy the commodity layer, build the part that is genuinely yours, and integrate them. If a subscription solves it, we will say so on the first call.
Tell us what the software has to do
Bring the process, the constraint or the half-working spreadsheet. We will tell you whether it should be built, bought or left alone — and if it should be built, what the first cycle looks like.
Scoped estimate before any build starts · Your code, your infrastructure, your data