A working agent demo does not answer the questions that appear in production.
What happens when retrieval pulls the wrong internal document? How much latency can a multi-step workflow tolerate? Can the team tell why an agent chose one tool over another? Will a prompt or model change break behavior that worked last week?
These are engineering problems, and they tend to appear after the prototype already looks convincing.
The Build Track at ODSC AI NYC is organized around that stage of the work. The two-day, workshop-driven AI summit takes place December 2–3 in New York City, with the Build Track aimed at AI engineers, ML engineers, data scientists, architects, and other technical practitioners working on agents, RAG systems, evaluations, observability, and production AI workflows.
For technical teams comparing AI and ML conferences, the distinction is useful. The Build Track is structured around working sessions and longer workshops where attendees can get into the implementation details behind production systems.
Build AI workflows that actually work. Join the ODSC AI NYC: AI for Work Summit in New York, December 2–3, for hands-on workshops focused on AI agents, workflow automation, evaluation, governance, and enterprise AI deployment. Explore the summit and register →
Production changes the engineering problem
An agent that works in a controlled test environment has several advantages. The data is known. The task is usually narrow. The person running the demo is watching closely enough to catch a strange result. Production removes those protections.
The agent encounters incomplete instructions, older documents, conflicting records, tool failures, and users who do not know what the system expected them to provide. A model update can also change behavior without producing the kind of obvious error a conventional test suite would catch.
The current Build Track agenda focuses on five areas that show up repeatedly in these systems: agent reliability, evaluation, observability, retrieval over permission-sensitive data, and choosing models according to cost, latency, quality, and risk.
That combination matters because those problems are connected. A retrieval error may look like a model failure. A slower model may improve one eval while making the complete workflow too expensive. An agent may produce the right final answer after taking a path that the team would not want repeated.
Agent observability needs to show what happened inside the run
Traditional application monitoring can tell a team whether an API request succeeded, how long it took, and whether the service threw an error. That is not enough for an agent.
Teams also need to see the retrieval steps, tool calls, model interactions, and decisions that produced the result. Without that trace, a successful HTTP request can hide a poor execution path.
Michael Levan, AI Architect and Forward Deployed Engineer at solo.io, will lead the advanced workshop Instrumenting an Agent for Observability. The session covers span design for reasoning steps, tool calls, and retrieval; OpenTelemetry GenAI conventions; context propagation to MCP servers; and the use of production traces to build evaluation sets. Attendees will also work backward through an agent failure trace to identify its cause.
That is a useful distinction for teams building their own monitoring stack. Service health and agent behavior need separate instrumentation.
ODSC recently covered the same issue in LLM Observability: Your Dashboard Is Green. Your Agent Is Wrong, which looks at failures that infrastructure metrics alone will not expose. Read it on opendatascience.com →
Evaluation has to run during development
An evaluation suite is much less useful if the team runs it only before launch.
Agent behavior can change when the model changes, but also when somebody edits a tool description, changes retrieval settings, modifies the system prompt, or adds another step to the workflow.
Those changes do not always cause a crash. The system can continue to run while making worse decisions.
Bruno Gonçalves’ four-hour Build workshop, You Can’t Ship What You Can’t Measure: Building a Production-Grade LLM Eval Harness, is built around this problem. Attendees will create an automated evaluation harness and turn it into a regression suite that can run in CI. The session combines deterministic answer matching, model-based grading, human spot checks, confidence intervals, and significance testing.
The useful part is the deployment model. Evals become part of the engineering workflow rather than a separate quality exercise.
ODSC’s Agentic Workflows Need Regression Tests covers the same problem from the perspective of prompt changes, tool schemas, retrieval behavior, and execution paths. Read it on opendatascience.com →
For a broader treatment of agent evaluation before release, AI Evals & AIOps: How Teams Actually Test Agents Before They Reach Users looks at building evaluation into the development pipeline rather than treating it as the last step. Read it on opendatascience.com →
Retrieval gets harder with company data
RAG demos usually start with a clean document collection. Company data rarely stays clean.
The same policy may exist in several versions. Permissions vary by team. Important context may live in a PDF, a CRM record, and an internal wiki at the same time. Some documents are relevant but should not be visible to the user making the request.
The Build Track includes retrieval over messy, permission-sensitive enterprise data as a specific production problem.
This requires more than tuning chunk size or swapping embedding models. Teams need to test whether the correct source was found, whether the requesting user had access to it, and whether the model used the retrieved evidence properly.
The problem becomes more difficult once an agent controls retrieval. It may rewrite the search query, decide to search again, or skip retrieval altogether.
From RAG to Agents: An Incremental Path to Agentic AI traces that progression from fixed retrieval through query rewriting and optional retrieval to a system where retrieval becomes a tool the agent decides when to call. Read it on opendatascience.com →
Cost and latency belong in the evaluation
A model can score higher on an eval and still be the wrong model for a particular production step.
If a workflow calls the model once, latency may be manageable. An agent that plans, retrieves, calls several tools, checks its result, and retries can multiply both response time and inference cost.
The Build program treats this as part of system design. Model selection is framed around cost, latency, quality, and risk, and the Day 2 advanced sessions include cost and latency routing at scale.
That suggests a different way to benchmark production systems. Measure the complete task, not only an isolated model call.
A cheaper model may work well for classification or routing while a more capable model handles the smaller number of steps that require it. The eval suite should tell the team whether that routing strategy changes task quality, while traces show what it does to latency and cost.
The Build Track is still being filled
Michael Levan’s observability workshop and Bruno Gonçalves’ evaluation workshop already have detailed agendas.
Thomas J. Fan, Member of Technical Staff at Modal, and Christine Long, Software Engineering Manager for ML Engineering at Meta Reality Labs, are also among the first confirmed instructors. Their individual session titles have not yet been published.
The preliminary program also places multi-agent orchestration, context engineering, eval suites in CI, and cost and latency routing among the advanced Build topics planned for Day 2.
The full program is scheduled for October 1, with additional sessions and instructors being added as they are confirmed.
What to test before moving an agent into production
A team preparing an agent for production can start with one workflow that already works reasonably well.
- Instrument the complete run. Capture retrieval, model calls, tool use, latency, errors, and token consumption. If the agent fails, the team should be able to reconstruct what happened without guessing.
- Build an eval set from the tasks the system is expected to perform. Add failed production cases as they appear. Run that set again when the prompt, model, retrieval configuration, or tool definitions change.
- Test retrieval separately from generation. If the final answer is wrong, determine whether the system retrieved poor evidence or failed to use good evidence.
- Set cost and latency limits for the complete task. An agent that eventually produces the correct answer after unnecessary retries still has a production problem.
These are the kinds of implementation questions the Build Track is designed around.
ODSC AI NYC Build Track
ODSC AI NYC takes place December 2–3, 2026, at Jay Conference Center Bryant Park in New York City. The summit includes 30+ sessions and workshops across four tracks, with workshop capacity limited and total event capacity listed at up to 450 seats.
For experienced technical practitioners, the Build Track focuses on agents, evals, RAG, observability, retrieval, routing, and production AI engineering. View the Build Track and preliminary program →
Early Bird pricing ends Friday, October 9 at midnight ET, with savings of up to $551 on the current event page.

