# Nina Labs

> Nina Labs is an applied AI engineering firm. We design, build and run AI agents, agentic workflows and AI software factories inside companies, connected to their systems and measured with evals.

- Page: https://www.ninalabs.ai/
- Publisher: Nina Labs AI LLC

## What we do

We build AI agents and run them in production. Nina Labs designs agents, agentic workflows and AI software factories, connects them to the systems a company already runs, and measures them with evals. Everything we build lives in the client's repositories and cloud.

## Services

- [AI agents in production](https://www.ninalabs.ai/services/agents/): We design, build and run AI agents that work inside your systems. They read, decide, call tools and hand the case to a person when a rule says so.
- [Agentic workflow automation](https://www.ninalabs.ai/services/workflows/): We rebuild back-office processes as workflows. Fixed steps stay as code, agents handle the steps that need judgment, and people approve the exceptions.
- [AI software factory](https://www.ninalabs.ai/services/software-factory/): We set up coding agents across your delivery pipeline, from spec to merge. Your engineers own the specs, the review gates and the releases.
- [Claude and OpenAI team enablement](https://www.ninalabs.ai/services/enablement/): We roll out Claude, ChatGPT, Claude Code and Codex to your teams, then teach each team to use them in its daily work, with admin setup, data rules and a library of tested workflows.
- [Legacy code modernization](https://www.ninalabs.ai/services/modernization/): We use coding agents to document, test and migrate legacy systems in small, verified steps, from COBOL and old Java to aging .NET and monoliths.
- [MCP and systems integration](https://www.ninalabs.ai/services/integration/): We connect agents to your systems through Model Context Protocol (MCP) servers and APIs, with real authentication, scoped permissions and audit logs.
- [Evals, security and AI governance](https://www.ninalabs.ai/services/evals-governance/): We measure whether your AI systems work, test how they fail and document them for auditors, with evals, red-team tests, tracing and compliance mapping.
- [AI strategy and operating model](https://www.ninalabs.ai/services/strategy/): We help leadership pick the AI work that pays off, decide what to build or buy, and set up the operating model that keeps AI systems owned and measured.

## Claude and ChatGPT, rolled out team by team

We roll out Claude, ChatGPT, Claude Code and Codex, connect them to your tools, and train each team on its own work. When a workflow is ready, we turn it into a production agent.

### Anthropic

We roll out Claude to whole companies and build production agents on the Claude Agent SDK. Engineering teams get Claude Code set up on their own repositories.

- Claude Enterprise setup: SSO, roles, data controls and connectors to your tools
- Projects and Agent Skills for each team's repeat tasks
- Claude Code for engineers: CLAUDE.md files, hooks, permissions, review rules and CI
- Production agents on the Claude Agent SDK or Claude Managed Agents
- Claude through the Claude API, Amazon Bedrock, Google Cloud or Microsoft Foundry

### OpenAI

We roll out ChatGPT Enterprise and Codex, train each team on its real work, and build agents on the OpenAI Agents SDK and the Responses API.

- ChatGPT Enterprise setup: SSO, workspace roles, data controls and connectors
- Shared projects and workflows for each team's repeat tasks
- Codex for engineers: cloud tasks, CLI, IDE extension, AGENTS.md files and code review
- Production agents on the OpenAI Agents SDK and the Responses API, including realtime voice
- OpenAI models through the OpenAI API, Microsoft Foundry or Amazon Bedrock

## How an AI software factory works

Coding agents do most of the implementation work. Engineers write and approve the specs, own the review gates and decide what ships.

1. Ticket: an engineer describes the change and the acceptance criteria.
2. Spec: an agent drafts the spec and the plan; an engineer approves them.
3. Implement: a coding agent changes the code in a sandbox and runs the tests.
4. Review: a review agent checks the change against the spec, conventions and security rules.
5. Verify: CI runs the tests, the evals and the static analysis.
6. Merge: an engineer reviews and merges; the release process ships it.

## How an engagement runs

### 01. Map (1–2 weeks)

We sit with the people who do the work, measure the current process and pick the tasks where an AI system can finish most cases end to end.

Outputs: Ranked use cases; Baseline metrics; Architecture and risk notes.

### 02. Build (4–8 weeks)

We build the system in your environment: the agent or workflow, the tools, the permissions and the eval suite. Your engineers pair with us from day one.

Outputs: Working system in your cloud; Eval suite from real cases; Tracing and dashboards.

### 03. Prove (2–4 weeks)

The system runs in shadow mode on live cases while people do the real work. We compare the results and fix what the evals and the traces show.

Outputs: Shadow-mode results; Go-live criteria; Runbook and on-call guide.

### 04. Run (Ongoing)

We turn it on in stages, watch quality and cost, and re-run the evals on every model or prompt change. Then we hand it over to your team.

Outputs: Staged rollout; Monthly quality and cost report; Handover and exit plan.

## What every system we ship includes

- **Evals:** A suite of real cases runs on every change to prompts, tools or models.
- **Tracing:** Every step, model call and tool call is recorded in OpenTelemetry format.
- **Approvals:** Actions that move money, change records or contact customers need a person until the data says otherwise.
- **Permissions:** Each tool has a written scope. Agents act with the user's rights or a narrow service identity.
- **Budgets:** Cost, token and latency limits per run, with alerts before they are reached.
- **Hosting:** Your cloud account or VPC by default. Open-weight models when data must stay on your hardware.
- **Models:** Chosen per task on measured quality, latency and cost. Switching is a configuration change.
- **Ownership:** Code, prompts, eval sets and infrastructure live in your repositories from the first day.
- **Handover:** A runbook, training for your team and a written exit plan in every engagement.

## Labs and platforms we work with

- **Frontier labs:** Anthropic Claude, OpenAI GPT, Google Gemini, Meta, Mistral AI, Perplexity, Cohere, Grok
- **Open-weight models:** DeepSeek, Z.ai GLM, Qwen, Kimi, MiniMax, Llama, Gemma, NVIDIA Nemotron
- **Agent platforms:** Claude Agent SDK, OpenAI Agents SDK, Google ADK, Vertex AI, Microsoft Foundry, Amazon Bedrock, LangGraph, LlamaIndex
- **Coding agents:** Claude Code, OpenAI Codex, GitHub Copilot, Cursor, Kiro, Google Antigravity, Windsurf
- **Protocols:** Model Context Protocol, Agent2Agent (A2A), AGENTS.md, Agent Skills
- **Clouds and runtimes:** AWS, Microsoft Azure, Google Cloud, NVIDIA NIM, Hugging Face, vLLM, Ollama, Your own hardware
- **Orchestration:** Temporal, n8n, AWS Step Functions, Azure Durable Functions
- **Observability and evals:** OpenTelemetry, Langfuse, LangSmith, Braintrust, Arize Phoenix, Promptfoo

Details: https://www.ninalabs.ai/platforms/

## Frequently asked questions

### What does Nina Labs do?

We are an applied AI engineering firm. We design, build and run AI agents, agentic workflows and AI software factories inside companies, connect them to the systems those companies already use, and measure them with evals.

### What is an AI software factory?

A delivery pipeline where coding agents do most of the implementation work. Engineers write and approve the specs, own the review gates and decide what ships. We set up the agents, the repository groundwork, the review and eval gates, and the metrics.

### Which models and platforms do you work with?

We work with every major lab: Anthropic (Claude), OpenAI (GPT, ChatGPT and Codex), Google (Gemini and Vertex AI), Meta, Mistral and Perplexity, and with open-weight models from DeepSeek, Z.ai, Qwen, Kimi and MiniMax. We run them on Amazon Bedrock, Microsoft Foundry, Google Cloud or your own hardware, pick the model per task on measured quality, latency and cost, and build on open protocols such as MCP and A2A so you can switch later.

### Where do the systems run, and who owns them?

In your cloud account by default. The code, prompts, eval sets and infrastructure definitions live in your repositories from the first day. You own all of it.

### How do you control what an agent can do?

Every tool has a written scope, agents act with the user's permissions or a narrow service identity, and actions that move money, change records or contact customers need a person's approval until the evals and production data show otherwise. Every run is traced.

### How long does a first project take?

The mapping phase takes one to two weeks. A first system usually reaches shadow mode four to eight weeks after that, and goes live in stages over the following weeks.

### Can you train our teams to use Claude and ChatGPT?

Yes. We set up Claude Enterprise, ChatGPT Enterprise or both, write the data rules with your security team, and run hands-on workshops for each team on its own tasks. Engineers get Claude Code and Codex set up on their repositories. The result is a library of tested workflows that each team owns.

### How do we start?

Send us a short note through the contact page. An engineer replies, usually with a few questions and times for a 30-minute call. If an AI system is not the right tool for the problem, we will say so.

## Contact

Talk to an engineer through the [contact form](https://www.ninalabs.ai/contact/) or by email at contact@ninalabs.ai.
