Skip to content
DesignBench

Generated by Grok 4.6 · DurkBench design benchmark

9m 11s start to finish

Everything below this bar is the page the model wrote. DurkBench did not change a word of it, and the company it describes is made up.

Skip to content

Lisbon studio · Est. 2023

We design the systems behind the systems.

Nine senior practitioners in Lisbon who architect agents, evaluation harnesses, and retrieval for companies that already know a demo is not a system.

  • 14

    Systems in production

  • 9

    Specialists in Lisbon

  • 6 mo

    Typical retainer

  • 0

    Abandoned pilots

01 / Services

What we actually build

We do not sell a platform. We design the four layers that make an AI system operable by your own team.

01

Agent architectures

We specify the graph: tools, memory, handoffs, and the human checkpoints that keep a multi-agent system from improvising.

You leave with a written architecture, not a notebook of prompts.

02

Evaluation harnesses

We build the tests that tell you whether the agent is getting better, not louder.

Golden sets, graders, and regression gates sit in CI so a model swap cannot silently undo a quarter of work.

03

Retrieval pipelines

We design the index, the chunking, the query rewrite, and the failure modes when the corpus is messy.

The point is citations you can defend, not a vector store you cannot explain.

04

Runtime operations

We wire traces, budgets, kill switches, and on-call paths around the model so production is a system, not a hope.

This is the plumbing that lets a nine-person team sleep.

02 / Method

A system before a sprint

Implementation starts after the system is specified. The sequence below is the engagement, not a sales metaphor.

  1. 1

    Week 0–2

    Diagnose

    Two weeks in the codebase and the call recordings. We write down what the system actually does, not what the slide deck claims.

  2. 2

    Week 2–3

    Specify

    A short architecture note: components, interfaces, evals, and the operational envelope. Nothing is built until this document survives a sceptical principal engineer.

  3. 3

    Week 3–10

    Instrument

    We ship the harness first. Agents, retrieval, and tools are then added against failing tests, the same way you would grow a compiler.

  4. 4

    Week 10–16

    Hand over

    Runbooks, traces, and a six-week pairing period with your team. We do not leave you holding a system only we can operate.

03 / Work

Three systems in production

Named clients, measured outcomes. Each engagement left a harness and a runbook with the in-house team.

18m → 4m

Median handling time

Meridian Bank

Wholesale operations

Problem

A 40-person ops team pasted between six systems to answer credit questions, and a vendor chatbot hallucinated on 31% of policy answers.

What we built

A tool-using agent over the policy corpus and core banking APIs, with a human checkpoint for exceptions above a stated limit.

Outcome

Median handling time fell from 18 minutes to 4, with citations on every answer and a measured 2.1% error rate.

12% → 67%

Accept-as-draft rate

Vespera Health

Clinical documentation

Problem

Draft notes from a speech model were unusable; clinicians spent longer editing than they had spent dictating.

What we built

A constrained generation pipeline with specialty schemas, retrieval over local guidelines, and an eval set of 1,200 de-identified notes.

Outcome

Accept-as-draft rate rose from 12% to 67% across three specialties in twelve weeks.

41% → 94%

First-response SLA

Atlas Freight

Exception handling

Problem

Delay and customs exceptions arrived as unstructured email; junior coordinators triaged by folklore.

What we built

A retrieval-and-routing agent that classified exceptions, pulled the shipment graph, and drafted the next action for a human dispatcher.

Outcome

First-response SLA recovered from 41% to 94%; the same desk absorbed a 30% volume increase without new headcount.

04 / Studio

Four of the nine

Four of the nine who sign the work. The others are embedded with clients this quarter.

  • IC

    Inês Carvalho

    Founding partner, systems

    Previously led applied ML at a Lisbon payments firm. She writes the architecture notes and refuses to ship an agent without a kill switch.

  • RA

    Rui Almeida

    Principal, evaluation

    Built the first model-eval platform at a European telco. He treats prompts as untrusted input and graders as the product.

  • MK

    Marta Kovács

    Principal, retrieval

    Spent six years on search at a news publisher. She is the person who asks what happens when the corpus is wrong.

  • TB

    Tomás Berger

    Partner, operations

    Ex-SRE from a global marketplace. He owns traces, budgets, on-call, and the boring diagrams that keep 3 a.m. quiet.

05 / FAQ

Questions we are asked first

Do you train foundation models?

No. We design the system the model lives in: tools, memory, retrieval, eval, and runtime. If a fine-tune is warranted we will say so and work with whoever trains it.

How small an engagement will you take?

A diagnostic is two weeks and a written architecture. We rarely take implementation under a three-month retainer; the failure mode of short work is theatre.

Will you work inside our stack?

Yes. We adopt your cloud, your repo, and your observability. We are not a platform vendor and we do not leave behind a private runtime you cannot hire for.

Where is the team based?

The studio is in Lisbon, at a workshop in Marvila. The nine of us work remotely with clients in Europe and the US, with overlap hours as a contract, not a slogan.

Who actually does the work?

Two principals on the problem, one partner on the operating envelope, and your engineers paired from week three. We do not staff a dozen juniors to look busy.

06 / Contact

Bring the system, not the pitch deck

Tell us what is in production, what is failing, and what you cannot measure. We reply within three working days with a yes, a no, or a sharper question.

Or write directly to studio@orrery.systems.