Инженер MLOps среднего/старшего уровня

в Caspian Innovation Center

О роли

Infrastructure & Platform Development

  • Design and implement scalable ML infrastructure on premises and cloud platforms
  • Build and maintain ML experimentation and production environments
  • Develop and manage container orchestration systems for ML workloads
  • Implement GPU resource management and optimization strategies
  • Design storage solutions for datasets, models, and artifacts


ML Pipeline & Automation

  • Create CI/CD pipelines for ML model training, validation, and deployment
  • Implement automated model retraining and versioning systems
  • Build orchestration workflows for data processing and model training
  • Develop automated testing frameworks for ML models and pipelines
  • Design and implement feature stores for feature engineering and reuse


Monitoring & Operations

  • Implement model monitoring systems for performance, drift, and data quality
  • Set up logging, alerting, and observability for ML systems
  • Establish model governance and compliance tracking
  • Create dashboards for model performance and infrastructure metrics
  • Develop incident response procedures for production ML systems


Collaboration & Best Practices

  • Partner with data scientists and AI engineers to productionize ML models
  • Establish MLOps best practices and standards across teams
  • Provide technical guidance on deployment architectures
  • Document processes, systems, and runbooks
  • Mentor junior engineers and data scientists on MLOps practices

Xüsusi tələblər

Education

  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)
  • Master's degree preferred but not required with sufficient practical experience

Experience

  • 2+ years working as ML/Software/DevOps Engineer
  • Proven track record of building production ML systems at scale
  • Experience supporting data science teams in enterprise environments

Technical Skills

  • Strong proficiency in Python and some experience with at least one low-level programming language (C/C++, Go, Rust, or similar)
  • Deep understanding of containerization (Docker, Kubernetes)
  • Hands-on experience with CI/CD tools (Jenkins/GitLab CI/GitHub Actions, or similar)
  • Knowledge of ML frameworks (TensorFlow/PyTorch/scikit-learn, or similar)
  • Experience with workflow orchestration (Airflow/Kubeflow/Prefect, or similar)
  • Hands-on experience with experiment tracking tools (MLflow/ClearML, or similar)
  • Experience with monitoring and observability tools (Prometheus, Grafana, ELK stack, or similar) for infrastructure and ML system monitoring


Core Competencies

  • Solid understanding of ML lifecycle and model development processes
  • Strong Linux/Unix systems administration skills
  • Experience with version control systems (Git) and branching strategies
  • Knowledge of networking, security, and compliance in cloud and on-prem environments
  • Understanding of distributed computing and parallel processing
  • Knowledge of microservices architecture and API design

Technical skills:

  • Experience with cloud platforms (Azure ML, AWS SageMaker, or GCP Vertex AI)
  • Experience with GitOps practices and tools (ArgoCD, Flux, GitLab with GitOps, or similar) for declarative infrastructure and ML pipeline management
  • Hands-on experience with hyperparameter optimization tools (Optuna/Ray Tune/Hyperopt/Katib, or similar)
  • Experience with distributed training frameworks
  • Experience with model serving frameworks (TensorFlow Serving/TorchServe/Triton/ MLServer, or similar
  • Experience with GPU optimization (CUDA/TensorRT/ONNX Runtime, or similar)
  • Knowledge of GPU allocation, sharing, management and profiling

LLM Ops:

  • Experience with LLM inference frameworks (vLLM/TGI/TensorRT-LLM, or similar)
  • Familiarity with agent orchestration frameworks (LangChain/LangGraph/LlamaIndex, or similar)
  • Experience with LLM optimization: quantization, KV cache management, continuous batching
  • Experience with prompt engineering and versioning tools (LangSmith/PromptLayer/Weights & Biases Prompts/Helicone, or similar)

Soft Skills:

  • Strong problem-solving and debugging abilities
  • Excellent communication skills with both technical and non-technical stakeholders
  • Ability to work independently and manage multiple priorities
  • Collaborative mindset with emphasis on enabling others
  • Adaptability to rapidly changing technology landscape
  • Pragmatic approach to balancing innovation with reliability
Müraciət etmək üçün: [email protected]

Infrastructure & Platform Development

  • Design and implement scalable ML infrastructure on premises and cloud platforms
  • Build and maintain ML experimentation and production environments
  • Develop and manage container orchestration systems for ML workloads
  • Implement GPU resource management and optimization strategies
  • Design storage solutions for datasets, models, and artifacts


ML Pipeline & Automation

  • Create CI/CD pipelines for ML model training, validation, and deployment
  • Implement automated model retraining and versioning systems
  • Build orchestration workflows for data processing and model training
  • Develop automated testing frameworks for ML models and pipelines
  • Design and implement feature stores for feature engineering and reuse


Monitoring & Operations

  • Implement model monitoring systems for performance, drift, and data quality
  • Set up logging, alerting, and observability for ML systems
  • Establish model governance and compliance tracking
  • Create dashboards for model performance and infrastructure metrics
  • Develop incident response procedures for production ML systems


Collaboration & Best Practices

  • Partner with data scientists and AI engineers to productionize ML models
  • Establish MLOps best practices and standards across teams
  • Provide technical guidance on deployment architectures
  • Document processes, systems, and runbooks
  • Mentor junior engineers and data scientists on MLOps practices

Xüsusi tələblər

Education

  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)
  • Master's degree preferred but not required with sufficient practical experience

Experience

  • 2+ years working as ML/Software/DevOps Engineer
  • Proven track record of building production ML systems at scale
  • Experience supporting data science teams in enterprise environments

Technical Skills

  • Strong proficiency in Python and some experience with at least one low-level programming language (C/C++, Go, Rust, or similar)
  • Deep understanding of containerization (Docker, Kubernetes)
  • Hands-on experience with CI/CD tools (Jenkins/GitLab CI/GitHub Actions, or similar)
  • Knowledge of ML frameworks (TensorFlow/PyTorch/scikit-learn, or similar)
  • Experience with workflow orchestration (Airflow/Kubeflow/Prefect, or similar)
  • Hands-on experience with experiment tracking tools (MLflow/ClearML, or similar)
  • Experience with monitoring and observability tools (Prometheus, Grafana, ELK stack, or similar) for infrastructure and ML system monitoring


Core Competencies

  • Solid understanding of ML lifecycle and model development processes
  • Strong Linux/Unix systems administration skills
  • Experience with version control systems (Git) and branching strategies
  • Knowledge of networking, security, and compliance in cloud and on-prem environments
  • Understanding of distributed computing and parallel processing
  • Knowledge of microservices architecture and API design

Technical skills:

  • Experience with cloud platforms (Azure ML, AWS SageMaker, or GCP Vertex AI)
  • Experience with GitOps practices and tools (ArgoCD, Flux, GitLab with GitOps, or similar) for declarative infrastructure and ML pipeline management
  • Hands-on experience with hyperparameter optimization tools (Optuna/Ray Tune/Hyperopt/Katib, or similar)
  • Experience with distributed training frameworks
  • Experience with model serving frameworks (TensorFlow Serving/TorchServe/Triton/ MLServer, or similar
  • Experience with GPU optimization (CUDA/TensorRT/ONNX Runtime, or similar)
  • Knowledge of GPU allocation, sharing, management and profiling

LLM Ops:

  • Experience with LLM inference frameworks (vLLM/TGI/TensorRT-LLM, or similar)
  • Familiarity with agent orchestration frameworks (LangChain/LangGraph/LlamaIndex, or similar)
  • Experience with LLM optimization: quantization, KV cache management, continuous batching
  • Experience with prompt engineering and versioning tools (LangSmith/PromptLayer/Weights & Biases Prompts/Helicone, or similar)

Soft Skills:

  • Strong problem-solving and debugging abilities
  • Excellent communication skills with both technical and non-technical stakeholders
  • Ability to work independently and manage multiple priorities
  • Collaborative mindset with emphasis on enabling others
  • Adaptability to rapidly changing technology landscape
  • Pragmatic approach to balancing innovation with reliability
Müraciət etmək üçün: [email protected]
Локация
Bakı
Опыт
Senior
Занятость
Полная занятость
Зарплата
Не указана
Опубликовано
7 августа 2026

О компании

Caspian Innovation Center
İnformasiya və kommunikasiya, Proqram təminatı və İT xidmətləri · Baku
Все вакансии Caspian Innovation Center

Ваш отклик

Создайте профиль на HRX

С ним вы узнаете решение работодателя, будете получать рекомендации по подходящим вакансиям и быстро на них откликаться.

Резюме *

Загрузите файлPDF до 8 МБ

Нет резюме? Соберите за 5 минут

Как работодатель может с вами связаться?

Нажимая «Откликнуться на вакансию», вы принимаете Условия использования и Политику конфиденциальности.

Похожие по навыкам

Все похожие вакансии

Вакансии, где совпадает больше всего навыков из этой позиции.

SOCAR Tech
JuniorBakı
З/п не указана

Создаёт и поддерживает инфраструктуру и инструменты для разработки, развертывания и мониторинга ML моделей.

AutomationAutomated TestingModel MonitoringDocumentation+10
Откликнуться
Все похожие вакансии
Не указана
Инженер MLOps среднего/старшего уровня
Откликнуться