Skip to content
DesignBench

Generated by Opus 4.8 · DurkBench design benchmark

1m 56s start to finish

Everything below this bar is the page the model wrote. DurkBench did not change a word of it, and the company it describes is made up.

Skip to content

AI systems consultancy · Lisbon

We design the systems behind the systems.

Nine senior engineers who build the agent architectures, evaluation harnesses, and retrieval pipelines your AI product quietly depends on.

senior engineers, no juniors on your bill
9senior engineers, no juniors on your bill
AI systems shipped to production
40+AI systems shipped to production
typical time to a working system
6 wkstypical time to a working system
engagements handed off, not handcuffed
100%engagements handed off, not handcuffed

What we do

Four things, done properly.

We stay narrow on purpose. These are the parts of an AI product that break under load, and the parts we know cold.

01

Agent architecture

We design multi-step agent systems that fail loudly and recover cleanly. Tool boundaries, memory, and hand-off logic, drawn before a single token is spent.

02

Evaluation harnesses

You cannot ship what you cannot measure. We build offline and online eval suites, regression gates, and dashboards your engineers will actually keep green.

03

Retrieval pipelines

Chunking, indexing, reranking, and the caching layer around them. We tune retrieval against your real queries, not a synthetic benchmark nobody believes.

04

Operational plumbing

Rate limits, fallbacks, cost ceilings, and tracing. The unglamorous scaffolding that keeps an AI product upright at three in the morning.

How we work

A short, legible engagement.

No open-ended retainers. Four steps, a fixed shape, and a deliberate exit.

  1. Map the terrain

    Two weeks inside your codebase and your data. We interview the people who own the problem and write down what everyone already half-knows.

  2. Draw the system

    A single architecture document: components, contracts, and the failure modes we expect. You approve it before we build anything.

  3. Build in the open

    Small pull requests, paired with your engineers, against a live evaluation suite. Nothing lands without a measured reason to trust it.

  4. Hand over the keys

    We leave runbooks, tests, and a team that no longer needs us. A clean exit is the whole point of hiring us.

Case studies

Systems in production, with numbers.

Client names are fictional; the shape of the work is exactly what we do.

Meridian Health

Clinical operations

−96%
Problem
A triage assistant hallucinated dosages and no one could tell which release made it worse.
Built
A retrieval pipeline over vetted guidelines plus a 400-case eval gate wired into CI.

Unsupported answers cut from 12% to 0.4%.

Kessler Freight

Logistics

4.5×
Problem
Support agents copied tracking data by hand across four internal systems.
Built
A tool-using agent with strict schemas, human approval on writes, and full tracing.

Median ticket handling time dropped from 9m to 2m.

Aperture Legal

Contract review

98.2%
Problem
Their first RAG prototype answered fluently and cited clauses that did not exist.
Built
Span-level grounding, a reranker tuned on their corpus, and a citation audit view.

Citation accuracy verified at 98.2% on a held-out set.

Who you work with

The people on the calls.

Remote from Lisbon, four of them likely on your project.

  • Inês Carvalho

    Principal, systems design

    Fifteen years turning vague AI ambitions into architectures that survive contact with production.

  • Tomás Rebelo

    Lead, evaluation

    Believes a benchmark you cannot reproduce is a rumour, not a result.

  • Sofia Marques

    Retrieval engineer

    Has opinions about chunk boundaries that have, twice, saved a launch.

  • Daniel Okonkwo

    Platform & operations

    Keeps the pagers quiet and the cost dashboards honest.

FAQ

Questions we get first.

If yours isn't here, ask it in the form below.

How large an engagement do you take?

Most run six to twelve weeks with two of us embedded. We turn down work we cannot staff properly; nine people means we say no often.

Do you write code or just advise?

We write code. We pair with your engineers in your repository and leave the system in their hands, not a slide deck on a shared drive.

Which models and vendors do you use?

Whichever your problem and constraints demand. We stay deliberately model-agnostic and design so that swapping a provider is a config change, not a rebuild.

Can you work with our compliance requirements?

Yes. We have shipped inside HIPAA and GDPR boundaries, with data-residency and audit constraints. Bring your security team to the first call.

What does it cost?

A fixed weekly rate, scoped after a short paid discovery. No surprise invoices and no reselling of infrastructure you could buy directly.

Start a project

Tell us what's breaking.

A short note is enough. We reply within two working days, and the first call costs nothing.

Prefer email? hello@orrery.systems