# AI agents in production

> We design, build and run AI agents that work inside your systems. They read, decide, call tools and hand the case to a person when a rule says so.

- Page: https://www.ninalabs.ai/services/agents/
- Publisher: Nina Labs AI LLC
- Updated: 2026-10-07

## Definition

An AI agent is a model that works in a loop with tools. It reads the input, picks the next step, calls a system, checks the result and stops when the task is done or a rule hands it to a person.

## What you get

- A use-case assessment with a measured baseline for volume, handling time and error rate
- One production agent with scoped tool access and a written permission model
- An eval suite built from real cases that runs before every release
- Tracing, cost and latency dashboards for every run
- A runbook, an on-call guide and a handover to the team that will own the agent

## Typical use cases

- **Support resolution:** Reads the ticket, checks the order and account systems, then resolves the case or drafts a reply for an agent to approve.
- **Accounts payable:** Matches invoices to purchase orders, flags price and quantity variances, and posts to the ERP after approval.
- **Sales operations:** Researches accounts, keeps the CRM current and drafts follow-ups that a rep reviews before they go out.
- **IT and HR service desk:** Handles access requests, resets and policy questions, and links every answer to the source document.
- **Voice agents:** Takes scheduling, intake and status calls on realtime speech models, and transfers the caller with a summary when needed.
- **Research and analysis:** Collects data from internal and public sources, runs the analysis and writes a cited brief for an analyst to check.

## What we build

We build agents for tasks with clear inputs, a known definition of done and a measurable cost today. Good first candidates are high-volume and repetitive, but they still need reading and judgment: a support queue, an invoice inbox, a service desk.

We do not start with a model or a framework. We start with a sample of real cases and a question: what share of them could an agent finish end to end, and what does each one cost now?

## How we build one

1. **Pick the task.** We look at volume, cost per case, the cost of an error and how often the rules change.
2. **Write the rules down.** Every tool gets a scope. Each action is either allowed, allowed after approval or not allowed.
3. **Build the evals first.** We collect real cases, remove personal data where needed and define what a correct result is for each one.
4. **Run in shadow mode.** The agent works on live cases while your team does the real work. We compare the results.
5. **Turn it on in stages.** By queue, by customer segment or by value threshold, with a fallback to the human queue at every stage.

## What runs in production

- Tracing for every step, tool call and model call, in OpenTelemetry format.
- Budgets for cost, tokens and latency per run, with alerts.
- A kill switch and a fallback path to your human queue.
- Versioned prompts, tools and model settings. The eval suite runs on each change.

## Models

We do not sell a model. We pick one per task based on measured quality, latency and cost, and we keep the option to switch. Many agents use a large model for planning and a smaller, faster model for routine steps. When data must stay on your own hardware, we use open-weight models.

## Stack we use

Claude Agent SDK, Claude Managed Agents, OpenAI Agents SDK, Microsoft Foundry Agent Service, Copilot Studio, Google ADK, Gemini Enterprise Agent Platform, Amazon Bedrock AgentCore, Strands Agents, Model Context Protocol (MCP)

## What we measure

- Share of cases the agent completes without a person
- Accuracy against the eval set and against human review
- Cost per completed case
- Time from request to resolution

## Frequently asked questions

### How long does the first agent take?

The assessment takes one to two weeks. Most first agents reach shadow mode six to eight weeks after that, depending on how many systems they need to reach.

### Where does the agent run?

In your cloud account by default. If your security team prefers a managed runtime, we use the one from your model or cloud vendor, such as Claude Managed Agents, Microsoft Foundry, Gemini Enterprise Agent Platform or Amazon Bedrock AgentCore.

### What happens when the agent gets something wrong?

Actions that move money, change records or contact customers go through an approval step until the evals and the production data show the agent can do them alone. Every run is traced, so we can find the cause, add the case to the eval set and fix it.

### Who owns the code and the prompts?

You do. The code, prompts, tool definitions, eval sets and infrastructure definitions live in your repositories from the first day.

## Related services

- [Agentic workflow automation](https://www.ninalabs.ai/services/workflows/)
- [MCP and systems integration](https://www.ninalabs.ai/services/integration/)

## Contact

Talk to an engineer through the [contact form](https://www.ninalabs.ai/contact/) or by email at contact@ninalabs.ai.
