Your brand on every pageBecome a sponsor →
citi logo

Generative AI Quality Engineer - Assistant Vice President

Citi

Pune Maharashtra India; DLF CYBERCITY 12B

Applied AI

Posted 10 days ago

vpunknown

Verified open on Oct 4, 2026 Ā· posted 9 days ago

Job Description

We are seeking a highly motivated and experienced AI Quality EngineerĀ to join our Retail and Wealth Risk EngineeringĀ team under the Enterprise Risk TechnologyĀ platform.Ā This role spans the full spectrum of modern AI quality engineering — fromĀ Agentic AI flow testingĀ andĀ RAG pipeline validationĀ toĀ AI safety, test automation, andĀ performance & reliability engineering.

You will be the quality pillar for complex autonomous AI systems, ensuring they areĀ safe, accurate, explainable, resilient, and production-readyĀ at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.

Responsibilities

AgenticĀ AI Testing

  • Design and executeĀ end-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
  • ValidateĀ agent reasoning, planning, and decision-making chainsĀ (e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
  • TestĀ tool-use correctness — ensuring agents invoke the right tools, with correct parameters, at the right time.
  • EvaluateĀ agent memory systemsĀ (short-term, long-term, episodic) for accuracy and context retention across sessions.
  • ValidateĀ agent handoff and delegation logicĀ in multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • TestĀ termination conditions, loop detection, andĀ infinite loop preventionĀ in autonomous agent loops.

RAG (Retrieval-Augmented Generation) Testing

  • Design comprehensive test strategies forĀ end-to-end RAG pipelines — covering ingestion, chunking, embedding, retrieval, reranking, and generation stages.
  • ValidateĀ retrieval accuracy and relevance — ensuring the correct context chunks are retrieved for a given query.
  • TestĀ embedding model qualityĀ and vector similarity thresholds across different document corpora.
  • EvaluateĀ faithfulness, groundedness, and answer relevanceĀ of generated responses using frameworks likeĀ RAGAS, TruLens, DeepEval.
  • TestĀ chunking strategiesĀ (fixed, semantic, hierarchical) for their impact on retrieval quality.
  • ValidateĀ context window management — ensuring retrieved context does not exceed token limits or degrade generation quality.
  • ConductĀ end-to-end regression testingĀ when the underlying knowledge base, embedding model, or LLM changes.
  • TestĀ multi-turn conversational RAGĀ for context coherence and citation accuracy across turns.

Test Automation

  • Build and maintainĀ automated test harnessesĀ for Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
  • DevelopĀ automated evaluation pipelinesĀ integrated into CI/CD workflows for continuous model and agent validation.
  • CreateĀ data validation and data quality frameworksĀ (using Great Expectations, Deequ, or custom tooling) for training, retrieval, and inference data.
  • BuildĀ prompt regression suitesĀ to detect behavioral drift across LLM versions or prompt changes.
  • ImplementĀ determinism and reproducibility testsĀ for stochastic LLM-based decisions.
  • AutomateĀ vector database validation — index integrity, embedding drift, and retrieval consistency checks.
AI Safety & Security Testing
  • ConductĀ red-teaming and adversarial testingĀ to uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
  • TestĀ output guardrails and content filtersĀ for unsafe, biased, toxic, or out-of-scope model behavior.
  • ValidateĀ privilege escalation controls — ensuring agents do not exceed permitted actions or access unauthorized resources.
  • PerformĀ data poisoning and backdoor attack simulationsĀ to assess model robustness.
  • Evaluate models forĀ bias, fairness, and discriminationĀ using frameworks such as AI Fairness 360 and Aequitas.
  • TestĀ PII leakage and data privacy controlsĀ in RAG and agent pipelines in accordance with GDPR, CCPA, and internal data governance policies.
  • Conduct security testing aligned with theĀ OWASP Top 10 for LLM Applications, including:
    • Prompt Injection (Direct & Indirect)
    • Insecure Output Handling
    • Training Data Poisoning
    • Insecure Plugin / Tool Design
    • Sensitive Information Disclosure
  • ValidateĀ constitutional AI constraints, RLHF-aligned behavior boundaries, and system prompt integrity.
  • Collaborate with cybersecurity teams onĀ AI-specific threat modelingĀ and vulnerability management.
  • MaintainĀ safety testing playbooksĀ and document red-team findings with severity ratings and remediation recommendations.

Performance & Reliability Testing

  • Define and executeĀ load, stress, soak, and spike testingĀ for AI-powered APIs, inference endpoints, and agent orchestration services.
  • Measure and optimizeĀ end-to-end latencyĀ across RAG and agentic pipelines — from query to final response.
  • BenchmarkĀ LLM inference throughputĀ (tokens/second) and identify bottlenecks across model serving infrastructure.
  • TestĀ auto-scaling behaviorĀ of AI services under variable load conditions.
  • ValidateĀ circuit breaker, retry, and fallback mechanismsĀ in agentic and RAG systems for graceful degradation.
  • TestĀ vector database performance — query latency, index build time, and retrieval accuracy under high concurrency.
  • ConductĀ cost efficiency analysis — measuring token consumption, API call costs, and infrastructure spend per agent task.
  • EstablishĀ SLOs (Service Level Objectives)Ā andĀ SLAsĀ for AI system availability, latency percentiles (P50, P95, P99), and error rates.
  • Collaborate with MLOps teams to set upĀ observability dashboards, monitoring alerts, and automated anomaly detection for production AI systems.
  • PerformĀ chaos engineering experimentsĀ to validate agent and RAG system resilience under infrastructure failures.

Domain Knowledge

  • Deep understanding ofĀ RAG architecture patterns — naive RAG, advanced RAG, modular RAG.
  • Solid grasp ofĀ agent design patterns: ReAct, Plan-and-Execute, Reflexion, MRKL, Mixture-of-Agents.
  • Familiarity withĀ AI safety and alignmentĀ principles (RLHF, Constitutional AI, guardrail layers).
  • Knowledge ofĀ token economics, context management, and LLM cost optimization.
  • Proficiency inĀ performance engineeringĀ methodologies for distributed AI systems.

Preferred Qualifications

  • Experience withĀ MCP (Model Context Protocol)Ā or similar agentic communication standards.
  • Exposure toĀ multi-modal agent testingĀ (agents handling text, images, code, documents).
  • Experience inĀ regulated industriesĀ (banking, finance, healthcare) with strict compliance requirements.
  • Familiarity withĀ chaos engineeringĀ tools (Chaos Monkey, Gremlin, LitmusChaos).

Education

  • Bachelor’s degree in Computer Science, Engineering, or a related field.

  • Master’s degree is a plus.

Experience
  • 8+ yearsĀ of experience in software or AI/ML quality engineering.
  • 3+ yearsĀ of hands-on experience withĀ RAG systems, or Agentic AI.
  • Proven experience buildingĀ automated test frameworksĀ for non-deterministic AI systems.
  • Strong background inĀ performance testingĀ andĀ AI safety/security assessments.

------------------------------------------------------

Job Family Group:

Technology

------------------------------------------------------

Job Family:

Technology Quality

------------------------------------------------------

Time Type:

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

Ā 

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Share this role:WhatsAppLinkedInX

Not ready to apply?

Get new Applied AI jobs in your inbox

Join 100+ AI professionals Ā· Weekly, free, unsubscribe anytime

Career context

This role is classified as Applied AI.

Preparing to interview? Read the Applied AI interview guide.

$195k - $255k is the middle 50% of disclosed salaries, measured from 425 live Applied AI postings on this board. Roughly two thirds of postings disclose nothing, so this describes the ones that do, not the whole market.

Share your AI salary

Anonymous. No login, no name, no email required. It takes about 20 seconds, and it is how this board builds real pay data for AI roles instead of guesses.

Before bonus and equity.

Hiring for a role like this?

Reach AI professionals browsing the board - your listing goes live instantly.

Post a job →

Related AI jobs

Join 100+ AI professionals. Get the week's new AI jobs in your inbox. Free, unsubscribe anytime.