Skip to main content
Duspat

Service

Agentic AI Workflows

Agentic workflows with explicit tool boundaries, human review points and audit trails — integrated into the systems where the work actually happens.

Who it is for

  • A team with a working LLM prototype and no route through security review to production
  • An operations leader with high-volume, judgement-heavy processes and a flat headcount
  • A Head of AI who needs a written control model before anyone will approve a deployment

Problems we solve

What tends to be going wrong.

Unbounded agents nobody will approve
An agent with broad tool access, no audit trail and no answer to the only question risk will ask: what can it do, and what happens when it is wrong?
Retrieval that leaks
A RAG system that answers confidently from documents the user was never entitled to open. The model is working exactly as built. That is the problem.
No way to tell whether it got worse
A prompt change ships on Friday. Nobody can say whether quality improved, degraded or stayed flat, because there is no evaluation harness and no regression suite.

How we help

What Duspat delivers.

01
Agent design and orchestration
Multi-step workflows with explicit tool boundaries — what an agent may call, what it may never call, what it must escalate, and where a human has to sign.
02
RAG and knowledge systems
Retrieval architecture over your own material, with entitlement-aware search so the model physically cannot surface what the reader is not cleared to see.
03
Human-in-the-loop workflows
Review points placed where the cost of being wrong is highest, designed so the reviewer sees enough context to actually judge — not a rubber stamp.
04
Evaluation and guardrails
An evaluation harness, regression tests, audit logging, rate limiting and cost controls. This is what turns a convincing demo into something you can sign off.

Typical deliverables

Things you can hold, review and hand to a team.

A short suitability and design engagement, then build-and-deploy. We will tell you early if a deterministic pipeline is simply the better answer.

  • Use-case assessment against a suitability model — including a recommendation not to proceed where that is the right answer
  • Agent workflow and tool-boundary design
  • Retrieval and knowledge-base architecture (RAG, embeddings, pgvector, Bedrock)
  • Entitlement-aware retrieval, enforced at the data layer rather than in the prompt
  • Human-in-the-loop review points and escalation paths
  • Evaluation harness, regression suite and quality baseline
  • Audit logging, rate limiting and cost controls
  • Production deployment and handover

Technology fit

What we build this on.

We are independent: no reseller agreements and no partner quotas. If your estate points a different way, we will say so.

Models & platforms
  • Claude
  • Amazon Bedrock
Retrieval
  • RAG
  • Embeddings
  • pgvector
  • PostgreSQL
Serving & orchestration
  • Lambda
  • API Gateway
  • Step Functions
  • ECS
Controls
  • IAM
  • CloudWatch
  • Evaluation harness
  • Audit logging

Governance & production readiness

Designed in, not added later.

The control model is designed before the model is chosen. That single ordering is the difference between an AI system that reaches production and one that stalls indefinitely at review.

How we run an engagement

  • Explicit tool boundaries — an allow-list, never a deny-list
  • Entitlement-aware retrieval enforced at the data layer
  • Audit logging of every tool call, retrieval and decision
  • Human review at the points where being wrong is expensive
  • Evaluation and regression testing before a prompt or model change ships
  • Rate limiting and cost controls, so a loop is an incident rather than an invoice

Example outcomes

Evidence, not testimonials.

Anonymised and sector-level. Client names are withheld by design — the constraint and the architecture are the parts that carry any information.

  • Life sciences

    A governed RAG system for sensitive research data

    The constraint. Retrieval over sensitive research material, where the model must never surface content the user is not entitled to see.

    The control model was designed before the model was chosen: entitlement-aware retrieval, IAM boundaries and audit logging built in from the first commit. The result was a proof of concept that could credibly become production, rather than a demonstration that quietly stalled at the security review.

    • Amazon Bedrock
    • pgvector
    • Lambda
    • API Gateway
    • IAM
    • CloudWatch

Questions

Common questions about agentic ai.

What is an agentic AI workflow?

A workflow where a model decides which tools to call and in what order, rather than following a fixed script. That flexibility is the value and the risk: an agent is only useful in production when its tool boundaries, review points and audit trail are explicit, so you can answer what it can do and what happens when it is wrong.

When is an agentic approach the wrong choice?

Whenever the steps are known in advance. If you can write the sequence down, a deterministic pipeline is cheaper, faster, testable and auditable, and an agent adds cost and non-determinism for nothing. We will tell you this early — it is usually the most valuable thing we say.

What is RAG, and when is it better than fine-tuning?

Retrieval-augmented generation retrieves your own content and gives it to the model at question time. It is almost always the right starting point: it keeps answers current, lets you enforce entitlements at the data layer, and shows its sources. Fine-tuning changes behaviour and tone, not knowledge, and is rarely the fix people expect.

How do you stop an AI system surfacing data a user should not see?

By enforcing entitlements at the data layer, not in the prompt. Retrieval is filtered by the requesting user’s permissions before anything reaches the model. A system that relies on instructions telling the model not to reveal something is not a control; it is a hope.

How do you evaluate an AI workflow before production?

With an evaluation harness and a regression suite built against real examples, so a prompt or model change can be measured rather than guessed at. Without one, nobody can say whether Friday’s change made the system better or worse — and eventually that uncertainty stops the rollout.

Review an agentic workflow before it goes live.

A first call is technical, not a sales call. Bring a problem, an architecture, or a stalled initiative.

30 minutes, no pitch deck.