Responsibilities
- Carrying out fine-tuning processes of Large Language Models (LLMs) tailored to various use cases;
- Preparation, cleaning, and optimization of high-quality datasets for public services and the Azerbaijani language;
- Objective evaluation of fine-tuning results and continuous improvement of model performance;
- Selecting the correct technical solution among various approaches (Fine-Tuning, RAG, Prompt Engineering, and Guardrails);
- Model training, optimization on GPU clusters, and preparation for the production environment;
- Ensuring experiments, models, and datasets are managed in a traceable and reproducible manner.
RequirementsBusiness and General Skills:
- Minimum 2 years of real work experience in the field of Machine Learning or LLMs;
- Analytical thinking and ability to systematically solve complex problems;
- Ability to adapt technical solutions to business needs;
- Effective collaboration with cross-functional teams and experience working with Agile methodologies;
- Ability to quickly analyze scientific papers and convert them into practical solutions;
- Ability to correctly assess what problems a fine-tuning approach can and cannot solve.
Technical Skills:
- High-level proficiency in the Python programming language;
- Experience working with the PyTorch framework;
- Deep understanding of Transformer architecture (Attention, RoPE, Tokenization, Context Window, KV Cache);
- Practical experience with Supervised Fine-Tuning (SFT), LoRA, QLoRA, DPO, and other Preference Tuning methods;
- Ability to work with the Hugging Face ecosystem (Transformers, PEFT, TRL, Datasets, Accelerate);
- Experience with Axolotl, LLaMA-Factory, Unsloth, or equivalent fine-tuning frameworks;
- Practical knowledge of Distributed GPU Training technologies (DeepSpeed, FSDP, NCCL);
- Management of GPU resources in a Slurm environment (sbatch, srun, multi-node job management, restart/failure handling);
- Experience working with H100, H200, or equivalent high-performance GPUs;
- Ability to apply BF16/FP16, Flash Attention, Gradient Checkpointing, and Quantization technologies;
- Dataset Engineering experience (Cleaning, Deduplication, Contamination Detection, Formatting, and Quality Scoring);
- Model fine-tuning experience for multilingual and low-resource languages;
- Tokenizer analysis (Vocabulary Coverage, Fertility, and Token Efficiency Analysis);
- Practical knowledge of LLM evaluation methods:
- Comparison of Base and Fine-tuned models;
- Preparation of Domain-specific Evaluation Datasets;
- LLM-as-a-Judge approach;
- Human Evaluation;
- Regression Testing;
- Measurement of Hallucination and Grounding metrics;
- Ability to correctly choose among Fine-Tuning, Retrieval-Augmented Generation (RAG), Prompt Engineering, and Guardrails approaches;
- Experiment tracking experience via MLflow or Weights & Biases (W&B);
- Model, dataset, and experiment lineage principles and reproducibility concepts;
- Experience deploying models to a production environment using the vLLM platform;
- Ability to work independently with Docker, Linux, and Git tools;
- Experience executing the entire lifecycle from a Hugging Face checkpoint to distributed fine-tuning, evaluation, merged/exported checkpoint, and production inference.
Preferred Qualifications:
- Continued Pre-training and Domain Adaptive Pre-training experience;
- Model Merging techniques;
- Long Context Tuning;
- Multimodal Model Fine-Tuning;
- Synthetic Data Generation;
- Knowledge Distillation;
- Speculative Decoding and other inference optimization techniques;
- Experience in LLM Safety, Guardrails, and Red Teaming;
- Experience developing AI solutions for public sector, legal, or regulated domains.