Johnson & Johnson
Lead AI Operations Engineer
- Ubicación
- Ubicaciones: 3
- Jornada
- Jornada completa
- Publicada
- ayer
La oferta
En Johnson & Johnson creemos que la salud lo es todo. Nuestra fuerza en la innovación en la atención médica nos permite construir un mundo en el que se eviten, traten y curen enfermedades complejas, en el que los tratamientos sean más inteligentes y menos invasivos, y las soluciones sean personales. A través de nuestra experiencia en Medicina Innovadora y MedTech, estamos en una posición única para innovar en todo el espectro de soluciones de atención médica de hoy para ofrecer los avances del futuro y afectar profundamente a la salud para la humanidad. Obtenga más información en jnj.com
Guiados por nuestro Credo, en Johnson & Johnson somos responsables de nuestros empleados que trabajan con nosotros en todo el mundo. Proporcionamos un entorno laboral inclusivo donde cada persona es tratada como individuo. En Johnson & Johnson, respetamos la diversidad y la dignidad de nuestros empleados y reconocemos sus méritos.
Función de trabajo:
Technology Product & Platform Management
Subfunción de trabajo:
Technical Product Management
Categoría de trabajo:
Scientific/Technology
Todas las ubicaciones de colocación de trabajo:
Lisbon, Portugal, Madrid, Spain, Milano, Italy
Descripción del trabajo:
We are recruiting for a Lead AI Operations Engineer based in Milan - Italy ; Madrid ; Spain or Lisbon ; Portugal
The AI Operations Engineer is responsible for shaping, designing, implementing, and continuously improving the enterprise capabilities required to operate AI applications and AI agents safely, reliably, transparently, and cost-effectively at scale. The role combines hands-on AI platform engineering with Site Reliability Engineering, DevSecOps, LLMOps, AgentOps, FinOps, security, and compliance practices.
The AI Operations Engineer builds reusable operational capabilities across observability, runtime controls, cost management, auditability, incident response, and production support. The role works closely with the Agent Factory, AI Engineering, Data Platforms, Cloud Infrastructure, Cybersecurity, Privacy, Risk, Quality, and Responsible AI stakeholders.
The role does not own the end-to-end lifecycle management of agents or AI products. Agent design, development, functional evaluation, release content, product evolution, and retirement decisions remain with the Agent Factory and the relevant AI product teams. The AI Operations Engineer provides the shared operational platform, telemetry, controls, and guardrails that enable those teams to run AI solutions in production.
Key Responsibilities
AI Observability & Production Reliability
- Design and implement end-to-end observability for AI applications and agents, including prompts, responses, model calls, tool calls, retrieval steps, decision paths, latency, failures, token consumption, and session context.
- Establish common telemetry and distributed tracing across agent workflows, APIs, data services, vector stores, model endpoints, and external tools.
- Build operational dashboards and alerts covering availability, latency, errors, reliability, quality signals, policy violations, consumption, and service health.
- Define service-level indicators, service-level objectives, error budgets, alert thresholds, and operational readiness criteria for production AI services.
- Enable trace-based debugging, incident reconstruction, and controlled session replay while protecting confidential or sensitive information in logs.
- Monitor retrieval quality, data freshness, model and prompt regressions, anomalous agent loops, degraded tool performance, and unexpected runtime behavior.
- Lead technical root-cause analysis for AI platform and runtime incidents and convert findings into preventive controls, automation, and engineering improvements.
LLMOps & AgentOps Platform Enablement
- Engineer reusable pipelines, templates, and controls for configuration, prompt, model, and agent-component versioning across environments.
- Implement automated technical gates for deployment readiness, including integration tests, regression checks, operational validation, security checks, and observability coverage.
- Enable controlled rollout patterns such as canary releases, feature flags, model or provider routing, fallback strategies, and technical rollback mechanisms.
- Provide common operational tooling that supports multiple models, frameworks, clouds, and agent patterns without creating a separate operating process for each solution.
- Integrate functional evaluation signals supplied by the Agent Factory or AI product teams into deployment gates and runtime monitoring, while functional quality ownership remains with those teams.
- Maintain reusable runbooks, reference implementations, engineering standards, and paved-road patterns for production operation and support.
AI FinOps & Consumption Efficiency
- Create transparent metering, allocation, and showback capabilities by product, agent, workflow, model, environment, and business unit where the required identifiers are available.
- Monitor token consumption, model utilization, repeated or runaway loops, retrieval overhead, infrastructure usage, and cost per successful transaction or workflow.
- Implement budgets, thresholds, anomaly alerts, and runtime guardrails to detect and contain unexpected consumption.
- Partner with Agent Factory and product teams to optimize model selection, routing, context size, caching, batching, retries, and tool usage while preserving agreed quality and compliance requirements.
- Define operational unit economics and provide evidence for capacity planning, optimization priorities, and platform investment decisions.
Security, Governance & Compliance Engineering
- Embed security, privacy, Responsible AI, and compliance controls into the shared AI runtime and operational toolchain in line with enterprise policies and approved risk frameworks.
- Implement identity, role-based access control, least privilege, managed identities, secrets management, and segregation of duties for agents, tools, services, and operators.
- Engineer runtime controls for prompt injection, jailbreak attempts, unauthorized tool use, excessive permissions, data leakage, unsafe execution paths, and anomalous access patterns.
- Design privacy-aware logging, retention, redaction, and access patterns for prompts, responses, memory, traces, and audit evidence.
- Provide auditable records linking versions, configurations, identities, actions, approvals, policy decisions, and operational outcomes.
- Automate policy checks and evidence collection where feasible, partnering with Cybersecurity, Privacy, Quality, Legal, Risk, and Responsible AI stakeholders for control definition and approval.
- Support threat modelling, security testing, incident response, remediation, and continuous control improvement for the AI platform.
Operational Service Management & Enablement
- Define operating processes for monitoring, support, incident management, problem management, change management, escalation, and service recovery.
- Create production-readiness checklists, service acceptance criteria, on-call runbooks, escalation paths, recovery procedures, and business-continuity requirements.
- Clarify operational handoffs and accountability across AI Operations, the Agent Factory, AI product teams, and underlying platform owners.
- Drive automation that reduces manual reconstruction of agent behavior and shortens detection, diagnosis, containment, and recovery times.
- Produce clear technical documentation and enable engineering teams to adopt approved operational patterns through practical coaching and reusable examples.
Stakeholder Management
- Act as the technical bridge between the Agent Factory, AI product teams, platform engineering, cloud infrastructure, data platforms, architecture, cybersecurity, privacy, compliance, and service management functions.
- Translate operational, risk, and control requirements into implementable technical capabilities and platform standards.
- Influence engineering teams to adopt common observability, reliability, cost, security, and compliance patterns.
Governance, Risk & Compliance
- Ensure shared AI Operations capabilities support applicable enterprise requirements for security, privacy, Responsible AI, auditability, and regulatory compliance.
- Maintain traceability of operational controls, exceptions, evidence, ownership, and remediation actions.
Metrics & Continuous Improvement
- Define and monitor KPIs for service reliability, observability coverage, incident performance, cost efficiency, control coverage, audit evidence, and platform adoption.
- Use operational evidence to prioritize automation, reliability improvements, cost optimization, and risk reduction.
- Measure the proportion of production AI services covered by approved telemetry, alerts, runbooks, cost controls, and security guardrails.
Qualifications
Required
- Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, or a related discipline, or equivalent relevant experience.
- 5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, DevSecOps, MLOps, or production software engineering, including responsibility for live services.
- Practical experience operating Machine Learning, Generative AI, or agentic systems in production, beyond proofs of concept or prompt engineering.
- Hands-on capability in Python, APIs, infrastructure as code, and CI/CD practices using technologies such as Terraform, GitHub Actions, or Azure DevOps.
- Experience with cloud-native platforms, containers, orchestration, and enterprise cloud services, with Azure experience strongly valued.
- Experience implementing observability using logs, metrics, distributed traces, dashboards, and alerts, ideally using OpenTelemetry or equivalent standards.
- Strong understanding of LLM application patterns including RAG, vector stores, model gateways, tool calling, agent memory, and multi-agent orchestration.
- Working knowledge of identity and access management, secrets management, secure logging, threat modelling, privacy by design, and audit controls.
- Ability to turn ambiguous operational and control requirements into reusable technical capabilities, standards, automation, and runbooks.
- Strong stakeholder management, technical communication, and documentation skills.
Preferred
- Experience working in regulated industries with formal security, privacy, quality, validation, risk, or audit requirements.
- Hands-on experience with Azure OpenAI, Azure AI Foundry, Azure Monitor, Application Insights, Microsoft Fabric, or comparable cloud services.
- Experience with AI observability or evaluation platforms such as Langfuse, LangSmith, Arize, MLflow, or equivalent technologies.
- Experience operating heterogeneous model providers, multi-cloud services, or multiple agent frameworks.
- Knowledge of FinOps practices, cloud cost allocation, budgeting, and unit-economics measurement for AI workloads.
- Experience defining service-level objectives, on-call models, incident management, and operational-readiness gates for enterprise platforms.
Habilidades requeridas:
Habilidades deseables:
Alineación de clientes, Análisis del panorama competitivo, Análisis de requisitos, Ciclo de vida de desarrollo del software (SDLC), Colaboración, Credibilidad técnica, Desarrollo ágil de productos, Desarrollo de productos, Estrategias de productos, Expertos en tecnología, Gestión de desarrollo de software, Gestión de la parte interesada, Interacción persona-máquina (HCI), Investigación y desarrollo, Mejoras de productos, Organización, Orientación, Pensamiento crítico, Pronóstico de la demanda, Razonamiento analítico, Redacción técnica
El intervalo de salario base anual previsto para esta posición es:
€48.600,00 - €77.395,00
Beneficios:
Además del salario base, ofrecemos los siguientes beneficios*: Una bonus variable anual con un objetivo o target preestablecido (% del salario fijo) en función del pay grade/ubicación, donde el importe a percibir se basa en el desempeño que hayan tenido durante el año natural anterior tanto los empleados como las empresas, o comisiones de ventas. Además, ofrecemos días de vacaciones, permiso parental por un mínimo de 12 semanas, permiso por duelo, permiso para cuidadores, permiso de voluntariado, reembolso de bienestar, programas de salud financiera, física y mental. Además, ofrecemos premios por aniversario de servicio y reconocimiento, y con sujeción a los términos de sus respectivos planes, los empleados, y los dependientes elegibles de algunas ubicaciones, pueden participar en varios planes de seguro. Para obtener más información, visite Employee benefits | Supporting well-being & career growth | Johnson & Johnson Careers.
*Lo anterior se facilita con fines exclusivamente informativos. Los importes y los beneficios reales pueden variar según la ubicación y están sujetos a cambios.