C
Not Specified Permanent

San Jose, California · USA job

AI LLMOps Engineering Technical Leader

CISCO Systems

San Jose, California

Job description

Overview

In this role you drive the operationalization of Agentic AI and MLOps within Cisco's CX platform to deliver secure, scalable production AI. You collaborate with cross-functional teams to shape reliable, high-performance agent architectures and ML pipelines. You will implement telemetry, monitor costs and model behavior, and guide architecture choices to balance latency, cost, and accuracy. This is an opportunity to influence the AI landscape at scale and contribute to a mission of robust customer experience.

Compensation / Benefits
  • medical, dental and vision insurance
  • 401(k) with company match
  • paid parental leave
  • bonuses subject to policy
  • restricted stock units (RSUs)
  • paid time off including holidays and wellness days
Responsibilities
  • Operationalize autonomous agent architectures into production-grade environments
  • Define and establish robust MLOps practices and ML pipelines
  • Design model serving architectures for experiments to production handoff
  • Develop telemetry tools to monitor token usage, compute costs, and reasoning paths
  • Troubleshoot and resolve infrastructure challenges for containerized GPU-based models
  • Guide architectural decisions and mentor peers to keep platforms reliable and scalable
Key requirements
  • Bachelor's degree and 8+ years of related experience, or Master's and 6+ years of related experience
  • Software engineering and DevOps experience with scalable infrastructure
  • Experience operationalizing Large Language Models (OpenAI, Anthropic, Llama) and frameworks like LangChain or LangSmith
  • Python and/or Java/J2EE
  • API development using FastAPI, Flask, Spring Boot or similar
  • CI/CD for ML and familiarity with model tracking and MLOps tools (ClearML, MLflow, Weights & Biases)
  • Docker and Kubernetes
  • Strong verbal and written communication
  • Leadership and mentorship of junior engineers
  • Ability to negotiate trade-offs (latency vs. accuracy vs. cost) with cross-functional partners
  • Agentic AI development and orchestration
  • AI/ML infrastructure on GPUs and multi-agent workflows
  • GPU compute scaling and latency optimization

Explore related USA jobs

Similar jobs you may like

Related roles with a similar title and location.

Same category Same location Same country

Process Control Engineer III

Calpine Operating Services Company

Middletown, California, United States · 95461

Same category Same location Same country

Physical Design Engineer - Static Timing Analysis, Annapurna Labs, Cloud Scale Machine Learning

Annapurna Labs (U.S.) Inc.

Cupertino, California, United States · 95014

Same category Same location Same country

ASIC Design Engineer, Cloud-Scale Machine Learning Acceleration team - Annapurna Labs

Annapurna Labs (U.S.) Inc.

Cupertino, California, United States · 95014

Same category Same location Same country

Process Engineer

Avantor

Carpinteria, California, United States · 93014

Same category Same location Same country

Building Maintenance Technician

US AMR-Jones Lang LaSalle Americas, Inc.

Seal Beach, California, United States · 90740

Same category Same location Same country

RF Design Engineer, In Building

Communication Technology Services Inc

Corona, California, United States · 92880

Same category Same location Same country

Structural Principal Engineer

Cannon Corp

Irvine, California, United States · 92619

Same category Same location Same country

Aviation Maintenance Technician

American Airlines

Los Angeles, California, United States · 90001

Need help finding a job?

Chat with our AI assistant.