At Bristol Myers Squibb, our employees often ask, “Who are you working for?”—a question that fuels collaboration, accountability, and urgency in our work. Our purpose-driven culture inspires us to discover, develop, and deliver innovative medicines to prevail over serious diseases. We offer uniquely interesting and meaningful work, opportunities for growth, and a supportive environment that values inclusion, wellbeing, flexibility, and comprehensive benefits. This is work that transforms the lives of patients, and the careers of those who do it.
Position Overview The Senior Specialist of AI Operations for AI Applications plays a critical role in supporting, maintaining, and optimizing AI applications for BMS’s PDS. This role combines technical expertise in IT operations, cloud engineering, automation, and AIOps to ensure AI applications run reliably, securely, and at scale. The ideal candidate is a hands-on engineer with strong problem-solving skills and experience supporting production machine learning or data-intensive systems.
Based on your function, department or individual position, you will have the opportunity to discuss with your Manager the option to work remotely up to 50% of the time, over a two-week period, with the flexibility to choose the days that align with your collaboration needs.
Key Responsibilities AI Application Support & Production Operations
• Provide operational support for production AI applications, ensuring availability, performance, and stability.
• Troubleshoot and resolve issues across data pipelines, model deployments, microservices, and integration layers.
• Implement monitoring, logging, and alerting solutions tailored to AI workload behavior.
• Maintain runbooks, operational documentation, and incident response procedures.
Automation & AIOps Engineering
• Deploy and manage containerized workloads using Kubernetes or similar orchestration frameworks.
• Support and enhance CI/CD and AI orchestration used for model training, testing, deployment, and monitoring.
• Automate operational tasks such as environment provisioning, configuration, scaling, and failover processes.
• Partner with AI engineers to ensure smooth handoff from development to production operations.
System Monitoring, Reliability & Performance
• Develop and maintain reliability engineering practices for AI systems, including SLO/SLI definitions and capacity planning.
• Conduct root cause analysis (RCA) for incidents and drive continuous improvement to prevent recurrence.
• Optimize system performance, cost efficiency, and resource utilization.
• Develop AI Native observability dashboards
Cross-Functional Collaboration
• Work closely with AI/ML teams, software engineering, data engineering, product management, and IT security.
Qualifications Required
• Bachelor’s degree in Computer Science, Information Technology, Engineering, or related discipline (or equivalent experience).
• 3+ years of experience in IT operations, site reliability engineering (SRE), cloud engineering, or DevOps.
• Hands-on experience with cloud environments (AWS), including compute, networking, IAM, and storage.
• Strong proficiency with Kubernetes, Docker, and containerized application architectures.
• Experience with automation tools and IaC frameworks (Terraform, CloudFormation, Ansible, etc.).
• Familiarity with LLM/ML tools and workflows (e.g., MLflow, Kubeflow, SageMaker, Vertex AI, Databricks).
• Proficient in scripting languages such as Python, Bash, or PowerShell.
• Strong troubleshooting skills across distributed systems, databases, APIs, and pipelines.
Preferred
• Experience supporting AI applications or generative AI workloads (LLMs, vector databases, embedding pipelines).
• Knowledge of observability platforms (Langsmith, Prometheus, Grafana, ELK/EFK, Datadog, New Relic, Splunk).
• Familiarity with ITIL practices, change management, and incident management frameworks.
• Background in highly regulated environments (healthcare, finance, government, etc.).
• Certifications in cloud (AWS/GCP/Azure), DevOps, or MLOps.
Key Competencies
• Technical Depth: Strong understanding of distributed systems, cloud architecture, and AI/ML operational workflows.
• Problem Solving: Ability to diagnose complex issues and design durable solutions.
• Automation Mindset: Focused on eliminating manual toil and improving operational efficiency.
• Collaboration: Comfortable working across data, engineering, security, and product teams.
• Reliability Focus: Dedicated to building robust, resilient, and well-monitored systems.
About the Role’s Impact This role is central to ensuring that AI applications operate smoothly in production, enabling the organization to confidently scale its AI initiatives. The Senior Engineer will help build and refine the technological foundation that supports AI innovation and operational excellence.
Why you should apply
• You will help patients in their fight against serious diseases
• You will be part of a company that encourages excellence and innovation, respects diversity, develops leaders and values its employees.
• You’ll get a competitive salary and a great benefits package including, but not only, an annual bonus, pension contribution, family health insurance, 27 days of annual leave , access to BMS Cruiserath on-site gym and life assurance
#LI-Hybrid
We hire for skills and capabilities, not just credentials – if this role excites you, but doesn’t perfectly match your resume, we encourage you to apply anyway.
How We Work Where you work matters – because collaboration, innovation and patient impact happen in many settings. Our roles are structured across four work models: site-essential, site-by-design, field-based and remote-by-design. The model assigned to this role is based on its core responsibilities. Learn more at https://careers.bms.com/ways-of-working.
Supporting People with Disabilities BMS is dedicated to ensuring that people with disabilities can excel through a transparent recruitment process, reasonable workplace accommodations/adjustments and ongoing support in their roles. Applicants can request a reasonable workplace accommodation/adjustment prior to accepting a job offer. If you require reasonable accommodations/adjustments in completing this application, or in any part of the recruitment process, direct your inquiries to adastaffingsupport@bms.com. Visit careers.bms.com/eeo-accessibility to access our complete Equal Employment Opportunity statement.
Candidate Rights BMS will consider qualified applicants with arrest and conviction records, pursuant to applicable laws in your area.
For roles based in Los Angeles County only: If you live in or expect to work from Los Angeles County if hired for this position, please visit this page for important additional information: https://careers.bms.com/california-residents/
Data Protection We will never request payments, financial information, or social security numbers during our application or recruitment process. Learn more about protecting yourself at https://careers.bms.com/fraud-protection.
Any data processed in connection with role applications will be treated in accordance with applicable data privacy policies and regulations. If this posting is missing required information required by local law or incorrect, contact BMS at TAEnablement@bms.com with the Job Title and Requisition number. Do not send application-related inquiries to this email. To check your application status, please login to your Candidate Home Account.
R1605220 : Sr Specialist - AI Operations