Description & Requirements
WHAT MAKES US A GREAT PLACE TO WORK
We are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times.
Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.
WHO YOU’LL WORK WITH
As the premier consulting partner for the private equity industry, Bain's PEG boasts a global practice that is over three times larger than any competitor. Our network of over 1,000 professionals supports private equity and institutional investor clients through every stage of the investment life cycle, from deal generation and due diligence to portfolio value creation and exit planning.
Bain & Company is developing a suite of cutting-edge data and software solutions designed to revolutionize how the private equity industry uses data for investment insights and decision-making.
The PEG Innovation team's mission is to create analytical solutions for Bain clients, teams, and the broader institutional investor space using proprietary software and data products. This includes the development, commercialization, and daily management of Bain's proprietary datasets, data, and software businesses.
WHERE YOU’LL FIT WITHIN THE TEAM
Senior ML Engineers build and operate the serving, deployment, and LLMOps infrastructure that carries models, prompts, and retrieval pipelines from prototype to governed production. You own the model and prompt lifecycle end-to-end (packaging, registry, promotion, staged rollout, and rollback) and you build the evaluation harnesses, inference services, and observability that keep production ML measurable and reliable. You partner with Data Scientists to productionize the models they build, with Data Engineers on feature and embedding pipelines, and with the Agent / AI squad to serve ML and retrieval outputs into agent workflows. You set the standard for how production ML systems are built and mentor mid-level engineers. This is a hands-on engineering role: models and pipelines that cannot be deployed, measured, and operated are not the goal.
WHAT YOU'LL DO
Core ML Systems Deployment, Serving, and Operations (80%)
- Build, deploy, and operate production inference and serving systems for models, embeddings, and re-rankers: request batching, concurrency, and throughput tuning against latency and cost SLAs.
- Own the model and prompt lifecycle in MLflow: packaging, model registry governance, promotion workflows, staged rollout behind feature flags, and clean rollback.
- Build and maintain LLMOps tooling: prompt and instruction versioning, model-gateway configuration (e.g., Portkey), inference orchestration, and response caching and cost controls.
- Build and operate production RAG and retrieval pipelines end-to-end: structure-aware chunking, contextual embedding, hybrid vector plus keyword retrieval, cross-encoder re-ranking, and context assembly.
- Design and maintain model and retrieval evaluation frameworks: golden datasets, metric definitions, LLM-as-judge with calibration, regression gates in CI, and production drift monitoring.
- Instrument production ML systems with structured logs, OpenTelemetry spans, and Prometheus metrics: token usage, latency percentiles, retrieval hit rates, drift, and hallucination monitoring, with dashboards and alerting.
- Collaborate with the Agent / AI squad to serve model and retrieval outputs as structured tool responses consumed by the Agent Gateway; partner with Data Engineers on feature and embedding pipelines.
- Drive production ML incident response to resolution; treat deployment, monitoring, and maintenance as part of delivery.
Other (20%):
- Set and enforce engineering standards for ML and serving code; contribute to repository conventions and raise the bar for production practices.
- Mentor mid-level ML Engineers and Data Scientists on production ML and LLMOps practice; conduct thorough code reviews and enforce standards on PRs.
- Use AI coding assistants to accelerate pipeline scaffolding, evaluation-harness development, and serving-config authoring; review all generated code against production standards before committing.
- Use LLMs to generate first-draft documentation, runbooks, and evaluation reports; validate and refine outputs before publishing.
ABOUT YOU
- Bachelor’s degree in Computer Science, Engineering, Machine Learning, Data Science, Statistics, or a related field (or equivalent practical experience).
- 6+ years of experience building and operating production ML systems, including model deployment, serving, and post-deployment monitoring.
- Demonstrated experience owning the model and prompt lifecycle end-to-end (packaging, registry, promotion, rollout, monitoring, and rollback).
- Demonstrated experience building and operating production RAG or retrieval systems end-to-end, from embedding and retrieval through re-ranking and evaluation.
- Experience collaborating cross-functionally with Data Science, Data Engineering, and Product teams to ship ML capabilities that solve real user problems.
- Strong Python for production ML; code written to production standards (testing, linting, typing).
- Experience working in a modern cloud ML platform environment (Databricks or AWS), including managed training / serving and governed model access.
- Demonstrated ability to mentor other engineers and raise engineering standards through code review and repository conventions.
ML engineering / LLMOps
- Strong Python: ML and serving code written to production engineering standards (type hints, Pydantic, pytest, Ruff, mypy strict).
- MLflow: experiment tracking, model registry, custom model flavours, promotion workflows, and model serving configuration.
- LLMOps tooling: prompt and instruction versioning, model gateways (e.g., Portkey), inference orchestration frameworks (LangChain, LlamaIndex, or equivalent), and response caching.
- Model serving and inference optimisation: request batching, concurrency and throughput tuning, latency budgeting, and awareness of quantisation and hardware trade-offs.
- RAG pipeline engineering: chunking strategies, contextual embedding, hybrid retrieval, cross-encoder re-ranking, and context assembly.
- Vector stores: pgvector, or dedicated vector databases; embedding pipeline design and index tuning at scale.
- Model evaluation and monitoring: golden datasets, metric definition, calibration, LLM-as-judge, regression gates in CI, and production drift monitoring.
- Feature Store integration: consuming point-in-time correct features (Databricks Feature Store or SageMaker Feature Store) in training and inference.
- Docker and Kubernetes: containerising training / inference workloads, writing Job and CronJob manifests, and understanding ephemeral workload patterns.
- Infrastructure familiarity: able to provision and review ML-serving infrastructure (IAM roles, model endpoints, GPU / CPU workloads) via Terraform without hand-holding.
Generative AI and agentic systems
- Builds and maintains inference and retrieval services that feed agent workflows as structured tool responses consumed by the Agent Gateway.
- Owns RAG serving quality: designs embedding and retrieval strategies and defines recall / precision and groundedness benchmarks.
- Uses LLM-as-judge patterns in evaluation pipelines where qualitative criteria are required.
- Uses production feedback and correction signals to design feedback-to-evaluation and feedback-to-training-data loops that close the loop between behaviour and model updates.
General
- Treats every ML system as a production system from the first commit: tests, observability, a runbook, and an SLA are not optional.
- Raises model quality and reliability issues proactively; does not wait for users to report degraded retrieval or drift.
- Uses AI tooling to move faster, but reviews all generated code and documentation critically before it enters the codebase.
- Strong communication skills; able to explain retrieval-quality, latency, and cost trade-offs to both engineers and non-technical stakeholders.
- This role follows a hybrid model, requiring in-office presence at least 1 day per week
U.S. COMPENSATION INFORMATION
Compensation for this role includes base salary, annual discretionary performance bonus, 401(k) plan with an annual employer contribution based on years of service and Bain’s best in class benefits package (details listed below).
Some local governments in the United States require a good-faith, reasonable salary range be included in job postings for open roles. The estimated annualized compensation for this role is as follows:
In Atlanta, the good-faith, reasonable annualized full-time salary range for this role is between $140,875 - $153,750
In Texas, the good-faith, reasonable annualized full-time salary range for this role is between $148,000 - $161,500
In Chicago, the good-faith, reasonable annualized full-time salary range for this role is between $155,125 - $169,250
Placement within these ranges will vary based on factors such as experience, education, training, and skill level.
Compensation also includes a discretionary annual performance bonus, 401(k) plan with employer contribution, and Bain’s best-in-class benefits—including full premium coverage for medical, dental, and vision, generous paid time off, and more.
Annual discretionary performance bonus
This role may also be eligible for other elements of discretionary compensation
4.5% 401(k) company contribution, which increases after 3 years of service and is 100% vested upon start date
Bain & Company's comprehensive benefits and wellness program is designed to help employees achieve personal independence, protection and stability in the areas most important to you and your family.
Bain pays 100% individual employee premiums for medical, dental and vision programs, offering one of the most comprehensive medical plans for employees without impacting your paycheck
Generous paid time off, including parental leave, sick leave and paid holidays
Fully vested 401(k) company contribution
Paid Life and Long-Term Disability insurance
Annual fitness reimbursements