Проводит тестирование мобильных и веб-приложений, требуется опыт автоматизации API тестов и наставничества
Старший инженер по контролю качества и оценке AI
О роли
About DataSpecta
DataSpecta is a Data & AI consulting and engineering company helping organizations transform data into measurable business value through advanced analytics, modern data platforms and enterprise AI solutions.
We design and deliver solutions across both on-premise and cloud environments, working closely with enterprise customers from discovery through production.
Alongside customer projects, we build our own AI products such as CogniSpecta (enterprise knowledge & cognitive search) and CodeSpecta (AI-powered software engineering).
We're looking for a Senior AI QA / Evaluation Engineer based in Baku, Azerbaijan, with strong test automation expertise and hands-on experience evaluating LLM and ML systems.
You will own quality assurance for production AI systems in a regulated enterprise environment, acting as the quality gate for every production release.
Location & Work Model
- Location: Baku, Azerbaijan
- Work model: [Hybrid / On-site]
What You Will Work On
- Evaluation harnesses and golden datasets for enterprise AI agents
- Evaluation of RAG assistants and report generation agents
- Adversarial and security testing for LLM applications
- Audit evidence and quality governance
Responsibilities
- Build versioned golden datasets and evaluation harnesses
- Gate production releases on evaluation baselines in CI
- Revalidate agents at each model version change
- Run quality control on recurring reports
- Run adversarial tests for prompt injection, data leakage and unsafe tool use
- Triage support and agent output incidents
- Maintain the audit evidence pack
Required Qualifications
- Based in Baku or willing to relocate
- 4+ years in QA or test engineering, including test automation ownership
- Proven experience evaluating LLM or ML outputs
- Strong Python skills for evaluation tooling
- Experience with formal defect management
- Fluent English
Technical Skills
- Evaluation frameworks: Ragas, DeepEval, promptfoo or equivalent
- Metrics: Groundedness, faithfulness, context precision/recall, answer correctness, tool call accuracy
- Statistics: Repeated sampling, bootstrapped confidence intervals
- Automation: Python, pytest, API testing, Playwright, CI
- Adversarial testing: OWASP Top 10 for LLMs, garak, PyRIT
- Observability: Langfuse, Arize Phoenix or equivalent
- Multilingual: Azerbaijani and mixed-language evaluation
Preferred Qualifications (a plus)
- Audit evidence experience in a regulated sector
- Model risk management frameworks
What We Value
- Curiosity about solving business problems with AI
- Comfort across business, data and technology
- End-to-end ownership
- Continuous learning
Follow us:
- LinkedIn: https://www.linkedin.com/company/innovance-consultancy
- LinkedIn: https://www.linkedin.com/company/dataspecta
- Instagram: https://www.instagram.com/innovanceconsultancy
- Instagram: https://www.instagram.com/dataspecta
About DataSpecta
DataSpecta is a Data & AI consulting and engineering company helping organizations transform data into measurable business value through advanced analytics, modern data platforms and enterprise AI solutions.
We design and deliver solutions across both on-premise and cloud environments, working closely with enterprise customers from discovery through production.
Alongside customer projects, we build our own AI products such as CogniSpecta (enterprise knowledge & cognitive search) and CodeSpecta (AI-powered software engineering).
We're looking for a Senior AI QA / Evaluation Engineer based in Baku, Azerbaijan, with strong test automation expertise and hands-on experience evaluating LLM and ML systems.
You will own quality assurance for production AI systems in a regulated enterprise environment, acting as the quality gate for every production release.
Location & Work Model
- Location: Baku, Azerbaijan
- Work model: [Hybrid / On-site]
What You Will Work On
- Evaluation harnesses and golden datasets for enterprise AI agents
- Evaluation of RAG assistants and report generation agents
- Adversarial and security testing for LLM applications
- Audit evidence and quality governance
Responsibilities
- Build versioned golden datasets and evaluation harnesses
- Gate production releases on evaluation baselines in CI
- Revalidate agents at each model version change
- Run quality control on recurring reports
- Run adversarial tests for prompt injection, data leakage and unsafe tool use
- Triage support and agent output incidents
- Maintain the audit evidence pack
Required Qualifications
- Based in Baku or willing to relocate
- 4+ years in QA or test engineering, including test automation ownership
- Proven experience evaluating LLM or ML outputs
- Strong Python skills for evaluation tooling
- Experience with formal defect management
- Fluent English
Technical Skills
- Evaluation frameworks: Ragas, DeepEval, promptfoo or equivalent
- Metrics: Groundedness, faithfulness, context precision/recall, answer correctness, tool call accuracy
- Statistics: Repeated sampling, bootstrapped confidence intervals
- Automation: Python, pytest, API testing, Playwright, CI
- Adversarial testing: OWASP Top 10 for LLMs, garak, PyRIT
- Observability: Langfuse, Arize Phoenix or equivalent
- Multilingual: Azerbaijani and mixed-language evaluation
Preferred Qualifications (a plus)
- Audit evidence experience in a regulated sector
- Model risk management frameworks
What We Value
- Curiosity about solving business problems with AI
- Comfort across business, data and technology
- End-to-end ownership
- Continuous learning
Follow us:
- LinkedIn: https://www.linkedin.com/company/innovance-consultancy
- LinkedIn: https://www.linkedin.com/company/dataspecta
- Instagram: https://www.instagram.com/innovanceconsultancy
- Instagram: https://www.instagram.com/dataspecta
Подходишь ли ты на эту вакансию?
Загрузи CV — за 10 секунд увидишь процент совпадения, свою цену на рынке и чего не хватает в резюме.
Зарплата на рынке
Работодатель сумму не назвал. По роли «QA-инженер» обычно платят 1 500–2 000 ₼, медиана — 1 500 ₼.
16 открытых вакансий, 13 зарплат: объявления, анкеты, «Проверь зарплату».
Проверить свою зарплату как «QA-инженер»О компании
Ваш отклик
Этот работодатель принимает отклики у себя на сайте. Перейдите на страницу вакансии и заполните форму там.
Перейти на страницу работодателяВопросы об этой вакансии
Рабочее место находится в городе Баку, в Бакинской экономической зоне.
Режим работы гибридный или в офисе.
Требуется минимум 4 года опыта в области обеспечения качества и тестирования.
Требуются Python, автоматизация тестирования, оценка LLM и ML, API тестирование и adversarial тесты.
Кнопка «Откликнуться» на этой странице ведёт на страницу отклика работодателя.
Перейти к форме откликаПохожие вакансии
Все похожие вакансииДругие открытые объявления по роли «QA-инженер» в городе Bakı.
Старший QA инженер обеспечивает качество программного обеспечения и требуется навык автоматизации тестирования.
Управляет процессами тестирования и обеспечения качества, требуются знания Java и Selenium.
Проводит тестирование критических бизнес-процессов и мобильной функциональности, требуется опыт QA от 2 лет.



