Required Skills: Snowflake, DataHub,
Job Description
We are looking for a Senior Technical Program Manager to own the operating model of AEGIS, a suite of production AI agents that answer real business questions for sales teams worldwide — forecasting, channel performance, and analytical insight drawn from governed enterprise sales data. Under the hood, AEGIS is a multi-agent system: an orchestration layer routes questions to specialized agents that plan, call tools, retrieve from governed data products, and compose answers grounded in enterprise metadata — business glossaries, metric definitions, and a knowledge graph that ties them together. Agent behavior is instrumented end to end with traces and scored continuously through an evaluation framework of golden sets, regression suites, and LLM-as-judge evaluators. The program spans agent development, data integration and governance, evaluation and quality, and the user-facing applications through which the agents are delivered, with regional sales and finance stakeholders feeding requirements from the field.
This is not a scrum-master or status-collection position. You will be accountable for how the program runs — the delivery cadence, release and quality governance, the decision forums, and the reporting executives trust — and you will inherit an active program mid-flight and be expected to stabilize and improve it. Because the product is a set of AI agents, program management here has a dimension most software programs don't: a prompt change, a retrieval tweak, a model version bump, or a shift in the underlying data can all change answers without touching application code. Evaluation is therefore a first-class release gate, regressions are measured statistically rather than caught by a failing unit test, and "done" is defined by eval results and production quality signals, not feature completion alone. You need to be fluent in that world and able to hold engineering and ML teams to it.
What You'll Do
Design and run the end-to-end delivery rhythm across engineering pods: planning, integrated release trains, review and retrospective forums, and a dependency map that is current and actually used.
Maintain a dated, dependency-aware roadmap with named owners, covering agent capabilities, data product onboarding, evaluation coverage, and platform infrastructure as a single plan of record.
Own release readiness and go/no-go decisions with engineering and cross functional leads, including rollback criteria, , and communication plans for user-visible behavior changes.
Own the intake-to-resolution pipeline for defects, data gaps, and enhancement requests: classification by layer (agent logic, retrieval, data quality, metadata, UI), routing to the owning team, SLAs, escalation paths, and closure verification.
Instrument the triage pipeline — resolution time, queue aging, reopen and rework rates, root-cause distribution by layer — and use the data to redesign handoffs where work stalls.
Lead root-cause analysis for user-visible quality incidents: read the traces, reproduce the failure, distinguish a bad retrieval from a bad plan from a bad data product, and drive remediation to verified closure.
Partner with operations to align evaluation and regression cycles to the release calendar so no agent, prompt, or model change reaches users without passing defined quality gates; own the gate definitions, thresholds by class of change, and the exception process.
Track eval coverage, regression trends, judge–human agreement, and production quality signals as program health metrics alongside delivery metrics.
Manage the data and metadata dependency backlog — new data products, metric definitions, glossary entries, and knowledge graph extensions — sequenced against agent capability delivery, and coordinate data contract changes and schema migrations so upstream changes don't silently degrade answers.
Own weekly program status and monthly executive reviews covering progress, quality, risk, and the decisions needed from leadership, written crisply and framed candidly.
Run the requirements front door for regional leads and field users with PM’s and data stewards capture, prioritize, translate into scoped engineering work with acceptance criteria and eval expectations, and close the loop back to requesters.
Must Have
10+ years of technical program management on complex software, data, or platform programs, including at least one multi-team program you ran end to end as the accountable owner.
Direct experience shipping LLM-based or ML products to production — you understand tool-calling agents, retrieval-augmented generation, prompt and model versioning, and why evaluation replaces deterministic testing as the release gate.
A track record of designing and instrumenting triage and quality operations, with measurable improvements in throughput, turnaround, or defect escape rate you can speak to specifically.
Technical depth sufficient to reason with engineers about APIs, data contracts and schemas, SQL and data warehouse concepts, logs and traces, and CI/CD pipelines — and to recognize when a proposed fix treats a symptom rather than the cause.
Executive presence: you have run reviews for senior leadership, delivered unwelcome news well, and driven decisions rather than reported on delays.
Exceptional written communication — status that reads in two minutes, specs engineers don't have to interpret, risk framing that is neither alarmist nor soft.
Demonstrated ability to enter an in-flight program and be effective within weeks.
Nice to Have
Hands-on experience with LLM observability and evaluation tooling (e.g., Langfuse, Arize Phoenix, Braintrust, LangSmith), including trace analysis, session-level scoring, LLM-as-judge design, and golden-set curation.
Experience taking an AI or data product from beta through general availability, including access management, user onboarding, and deprecation of legacy interfaces.
Working knowledge of modern data platforms — cloud data warehouses (e.g., Snowflake), data catalogs and metadata platforms (e.g., DataHub), graph databases, and data product or data mesh patterns.
Familiarity with knowledge graphs, ontologies, business glossaries, and semantic layers, and how they ground agent answers in governed definitions.
Domain literacy in sales operations, finance, or business intelligence — sell-in, sell-through, forecast accuracy, channel performance — and the ability to sanity-check whether an agent's answer is plausible.
Experience in a large, matrixed enterprise where influence and clarity matter more than authority.
Success in the First 90 Days
The architecture (agents, tools, retrieval, data products, eval pipeline), the teams, the forums, and the specific points where work stalls are mapped, with an explicit assessment of what to keep, fix, and retire in the current operating model.
An integrated cadence is running across pods with a single roadmap of record that leadership plans against.
Triage queues have defined SLAs, owners, layer-level classification, and weekly reporting; resolution time and cross-team rework are measurably down, with the metrics to show it.
Evaluation cycles are formally attached to release gates with documented thresholds, and exceptions are visible.
User feedback flows into the backlog through a repeatable path, requesters can see what happened to their input, and the operating model is documented well enough to hand to a permanent owner without loss of rigor.