Старший инженер по контролю качества и оценке AI

в dataSpecta

О роли

About DataSpecta

DataSpecta is a Data & AI consulting and engineering company helping organizations transform data into measurable business value through advanced analytics, modern data platforms and enterprise AI solutions.

We design and deliver solutions across both on-premise and cloud environments, working closely with enterprise customers from discovery through production.

Alongside customer projects, we build our own AI products such as CogniSpecta (enterprise knowledge & cognitive search) and CodeSpecta (AI-powered software engineering).

We're looking for a Senior AI QA / Evaluation Engineer based in Baku, Azerbaijan, with strong test automation expertise and hands-on experience evaluating LLM and ML systems.

You will own quality assurance for production AI systems in a regulated enterprise environment, acting as the quality gate for every production release.

Location & Work Model

  • Location: Baku, Azerbaijan
  • Work model: [Hybrid / On-site]

What You Will Work On

  • Evaluation harnesses and golden datasets for enterprise AI agents
  • Evaluation of RAG assistants and report generation agents
  • Adversarial and security testing for LLM applications
  • Audit evidence and quality governance

Responsibilities

  • Build versioned golden datasets and evaluation harnesses
  • Gate production releases on evaluation baselines in CI
  • Revalidate agents at each model version change
  • Run quality control on recurring reports
  • Run adversarial tests for prompt injection, data leakage and unsafe tool use
  • Triage support and agent output incidents
  • Maintain the audit evidence pack

Required Qualifications

  • Based in Baku or willing to relocate
  • 4+ years in QA or test engineering, including test automation ownership
  • Proven experience evaluating LLM or ML outputs
  • Strong Python skills for evaluation tooling
  • Experience with formal defect management
  • Fluent English

Technical Skills

  • Evaluation frameworks: Ragas, DeepEval, promptfoo or equivalent
  • Metrics: Groundedness, faithfulness, context precision/recall, answer correctness, tool call accuracy
  • Statistics: Repeated sampling, bootstrapped confidence intervals
  • Automation: Python, pytest, API testing, Playwright, CI
  • Adversarial testing: OWASP Top 10 for LLMs, garak, PyRIT
  • Observability: Langfuse, Arize Phoenix or equivalent
  • Multilingual: Azerbaijani and mixed-language evaluation

Preferred Qualifications (a plus)

  • Audit evidence experience in a regulated sector
  • Model risk management frameworks

What We Value

  • Curiosity about solving business problems with AI
  • Comfort across business, data and technology
  • End-to-end ownership
  • Continuous learning

Follow us:

  • LinkedIn: https://www.linkedin.com/company/innovance-consultancy
  • LinkedIn: https://www.linkedin.com/company/dataspecta
  • Instagram: https://www.instagram.com/innovanceconsultancy
  • Instagram: https://www.instagram.com/dataspecta

About DataSpecta

DataSpecta is a Data & AI consulting and engineering company helping organizations transform data into measurable business value through advanced analytics, modern data platforms and enterprise AI solutions.

We design and deliver solutions across both on-premise and cloud environments, working closely with enterprise customers from discovery through production.

Alongside customer projects, we build our own AI products such as CogniSpecta (enterprise knowledge & cognitive search) and CodeSpecta (AI-powered software engineering).

We're looking for a Senior AI QA / Evaluation Engineer based in Baku, Azerbaijan, with strong test automation expertise and hands-on experience evaluating LLM and ML systems.

You will own quality assurance for production AI systems in a regulated enterprise environment, acting as the quality gate for every production release.

Location & Work Model

  • Location: Baku, Azerbaijan
  • Work model: [Hybrid / On-site]

What You Will Work On

  • Evaluation harnesses and golden datasets for enterprise AI agents
  • Evaluation of RAG assistants and report generation agents
  • Adversarial and security testing for LLM applications
  • Audit evidence and quality governance

Responsibilities

  • Build versioned golden datasets and evaluation harnesses
  • Gate production releases on evaluation baselines in CI
  • Revalidate agents at each model version change
  • Run quality control on recurring reports
  • Run adversarial tests for prompt injection, data leakage and unsafe tool use
  • Triage support and agent output incidents
  • Maintain the audit evidence pack

Required Qualifications

  • Based in Baku or willing to relocate
  • 4+ years in QA or test engineering, including test automation ownership
  • Proven experience evaluating LLM or ML outputs
  • Strong Python skills for evaluation tooling
  • Experience with formal defect management
  • Fluent English

Technical Skills

  • Evaluation frameworks: Ragas, DeepEval, promptfoo or equivalent
  • Metrics: Groundedness, faithfulness, context precision/recall, answer correctness, tool call accuracy
  • Statistics: Repeated sampling, bootstrapped confidence intervals
  • Automation: Python, pytest, API testing, Playwright, CI
  • Adversarial testing: OWASP Top 10 for LLMs, garak, PyRIT
  • Observability: Langfuse, Arize Phoenix or equivalent
  • Multilingual: Azerbaijani and mixed-language evaluation

Preferred Qualifications (a plus)

  • Audit evidence experience in a regulated sector
  • Model risk management frameworks

What We Value

  • Curiosity about solving business problems with AI
  • Comfort across business, data and technology
  • End-to-end ownership
  • Continuous learning

Follow us:

  • LinkedIn: https://www.linkedin.com/company/innovance-consultancy
  • LinkedIn: https://www.linkedin.com/company/dataspecta
  • Instagram: https://www.instagram.com/innovanceconsultancy
  • Instagram: https://www.instagram.com/dataspecta

Подходишь ли ты на эту вакансию?

Загрузи CV — за 10 секунд увидишь процент совпадения, свою цену на рынке и чего не хватает в резюме.

Проверить CV
Локация
Bakı, İstanbul
Опыт
4+ лет
Занятость
Полная занятость
Зарплата
Не указана
Опубликовано
1 октября 2026
Языки
English
Просмотры
20

Зарплата на рынке

Работодатель сумму не назвал. По роли «QA-инженер» обычно платят 1 500–2 000 ₼, медиана — 1 500 ₼.

обычно платят столькоредко

16 открытых вакансий, 13 зарплат: объявления, анкеты, «Проверь зарплату».

Проверить свою зарплату как «QA-инженер»

О компании

dataSpecta
IT Services and IT Consulting · 11-50 · Bakı
Все вакансии dataSpecta

Ваш отклик

Этот работодатель принимает отклики у себя на сайте. Перейдите на страницу вакансии и заполните форму там.

Перейти на страницу работодателя

Вопросы об этой вакансии

Рабочее место находится в городе Баку, в Бакинской экономической зоне.

Режим работы гибридный или в офисе.

Требуется минимум 4 года опыта в области обеспечения качества и тестирования.

Требуются Python, автоматизация тестирования, оценка LLM и ML, API тестирование и adversarial тесты.

Кнопка «Откликнуться» на этой странице ведёт на страницу отклика работодателя.

Перейти к форме отклика

Похожие вакансии

Все похожие вакансии

Другие открытые объявления по роли «QA-инженер» в городе Bakı.

Predsol
Senior · 2+ летBakı
З/п не указана

Проводит тестирование критических бизнес-процессов и мобильной функциональности, требуется опыт QA от 2 лет.

Manual TestingTest DocumentationPlaywrightAutomation Testing+8
ОткликнутьсяНапрямую работодателюНапрямую работодателю
Все похожие вакансии
Не указана
Старший инженер по контролю качества и оценке AI
Откликнуться