Vice President AI Machine Learning Software positions focus on delivering results in their domain. This page aggregates open Vice President AI Machine Learning Software roles and what employers typically expect.
At BNY, our culture allows us to run our company better and enables employees’ growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the world’s investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide. Recognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance – and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary. **Role Overview** We are seeking a senior‑level engineer to design, build, and operate **production**‑**grade GenAI and Retrieval**‑**Augmented Generation (RAG) platforms** at scale. This role focuses on industrializing LLM‑based systems with strong **guardrails, observability, evaluation frameworks, and operational rigor**, ensuring reliability, safety, and cost efficiency across the full AI lifecycle. This role is located in Jersey City, NJ. **Key Responsibilities** - Design and build **production**‑**ready RAG pipelines**, including retrieval, ranking, prompt orchestration, and response generation, with comprehensive **guardrails, tracing, and observability**. - Implement **offline and online evaluation frameworks** for prompts, models, and datasets, including quality, safety, latency, and cost metrics. - Own **end**‑**to**‑**end lifecycle management** for GenAI systems, covering prompt versions, model versions, datasets, and configurations. - Establish and maintain **CI/CD pipelines** for prompts, models, and data, enabling safe, repeatable, and auditable releases. - Implement **cost and performance monitoring**, including token usage, inference latency, throughput, and spend optimization. - Build and enforce **safety mechanisms**, such as content filtering, policy enforcement, red‑teaming feedback loops, and abuse detection. - Define and operationalize **incident management workflows**, including alerting, triage, rollback mechanisms, and post‑incident analysis. - Partner closely with product, platform, and governance teams to ensure GenAI solutions meet **enterprise reliability, security, and compliance standards.** - Mentor engineers and influence best practices for building **scalable, trustworthy AI systems**. **What Success Looks Like** - GenAI systems that are **observable, measurable, and resilient**, not “black boxes.” - Safe and cost‑efficient RAG pipelines running reliably in production. - Fast iteration cycles with strong controls, enabling teams to ship GenAI features with confidence. **Required Qualifications** - Advanced degree in STEM engineering degree, or equivalent work experience with experience preferred in related fields. 7-9 years of related experience required; experience in the securities or financial services industry is a plus - Strong experience building and operating **production ML or GenAI systems** in enterprise environments. - Deep hands‑on expertise with **LLM orchestration frameworks**, such as **LangChain** and/or **LlamaIndex**. - Experience with **model registries and experiment tracking**, such as **MLflow** or equivalent. - Solid understanding of **Kubernetes**‑**based deployments** and cloud‑native architectures. - Familiarity with **feature stores**, data pipelines, and retriever/index lifecycle management. - Proven experience implementing **telemetry, logging, metrics, and distributed tracing** for ML/AI workloads. - Strong knowledge of **CI/CD practices** for ML, GenAI, and data‑driven systems. **Preferred Qualifications** - Experience operating **LLM systems at scale**, including multi‑model or multi‑provider strategies. - Exposure to **AI safety, governance, and compliance frameworks** in regulated environments. - Background in **SRE, platform engineering, or MLOps**, with a reliability‑firs…