open_to_work: AI QA · Test Automation

Sujitha A
AI Quality Engineer.

Senior AI Test Automation Engineer

I make AI systems measurable. 10 years in QA across Insurance, FinTech, SAP ERP and Salesforce, now evaluating LLMs, RAG pipelines and AI agents with Python, Playwright, DeepEval and Ragas.

sap-ai-agent — eval
$ pytest evals/sap_agent --golden-set 150
collected 150 scenarios · Order-to-Cash, Procure-to-Pay
PASS task_success 78% → 92%
PASS self_heal_rate ~60%
PASS false_failures 15% → 5%
PASS invalid_actions -40%
PASS regression_cycle 5d → 1.5d
══ quality gate: PASSED · ready to release ══
$
experience10 yrsQA & test automation
scenarios120+SAP flows automated by an AI agent
regression~70%faster SAP regression cycle
task_success~92%agent score on golden dataset
01 / about

From enterprise QA to AI quality

I started in functional and API testing for a FinTech lending app, then led QA on Salesforce CLM at IBM and Guidewire PolicyCenter at Capgemini, where I also tested UiPath RPA workflows across web, PDF, Excel and OCR.

Most recently I built a Python-based AI test agent for SAP GUI. It plans, runs and verifies end-to-end ERP scenarios, heals broken locators after SAP updates, and is measured against a golden dataset rather than judged by feel.

That is how I approach AI testing: define what "good" means, turn it into a metric, and let the number decide whether something ships.

{
  "role": "Senior AI Test Automation Engineer",
  "experience": "10 years",
  "domains": ["Insurance", "FinTech", "SAP ERP", "Salesforce"],
  "core": ["Python", "JavaScript", "TypeScript", "Playwright", "Pytest"],
  "ai_eval": ["DeepEval", "Ragas", "Promptfoo"],
  "location": "Hyderabad, India"
}
02 / stack

What I work with

# ai_llm_testing

LLM EvaluationRAG TestingAgentic AI Testing Hallucination DetectionGroundedness & FaithfulnessPrompt Injection Guardrail TestingLLM-as-a-JudgeGolden Datasets Bias & ToxicityMulti-turn TestingHuman-in-the-loop

# ai_frameworks

DeepEvalRagasPromptfoo LangChainLangGraphMCP OpenAI APIClaudeGeminiGitHub Copilot

# automation

PythonJavaScriptTypeScriptPlaywrightPytest SAP GUI ScriptingUiPathREST Assured PostmanSwaggerSQL

# delivery

CI/CDGitAzure DevOpsJira ZephyrAgile / ScrumSAFeTest Strategy
03 / approach

How I test AI systems

AI output is non-deterministic, so "it looked right" is not a test result. Every project runs through the same pipeline.

stage: define

Define quality

Agree with product owners what a correct, safe answer looks like and what must never happen.

stage: dataset

Build a golden set

Real questions and scenarios with expected answers, sources, edge cases and adversarial prompts.

stage: measure

Measure

Faithfulness, answer relevancy, retrieval recall, task success and tool-call accuracy, run in CI on every change.

stage: gate

Decide

Thresholds gate releases. Failures are triaged to retrieval, prompt, model, data or environment.

04 / projects

Selected work

rag_eval/RAG testing

Enterprise Knowledge Assistant

Evaluated a RAG assistant against a ground-truth question set to separate retrieval failures from generation failures.

  • Retrieval accuracy, citation correctness, chunk quality and answer faithfulness.
  • Adversarial and out-of-scope prompts to test fallbacks, guardrails, latency and token usage.
RagasDeepEvalLangChainOpenAI APIPytest
llm_eval/LLM evaluation

GenAI Chat Assistant Testing

Functional and regression testing for an LLM chat assistant, compared across OpenAI and Claude models.

  • Accuracy, hallucination, groundedness, consistency across prompt variations and multi-turn context.
  • Automated evaluations with structured failure reports for every run.
PromptfooDeepEvalOpenAI APIClaudePython
agentic_tests/agentic AI

Multi-Agent Workflow Testing

Tested a planner, executor and reviewer agent workflow with structured pass/fail criteria.

  • Planning accuracy, tool invocation, retries, memory and exception recovery.
  • Guardrail triggers and end-to-end behaviour, including browser actions via Playwright.
LangGraphMCPPlaywrightPython
uipath_rpa/Capgemini

UiPath Automation for P&C Insurance

RPA test automation for Farmers Insurance on Guidewire PolicyCenter.

  • Workflows across web, PDF, Excel, API and OCR, run through Orchestrator and Test Manager.
  • End-to-end traceability of tests and defects with Zephyr and Jira.
UiPathOrchestratorTest ManagerJira
05 / experience

git log --career

HEAD → ai-qualitySep 2025 – Jul 2026

Senior Automation Engineer

C Ahead Info Technologies · Client: Hitachi Digital Services · Toyota Services SAP GUI
  • Built and evaluated an AI test agent for SAP GUI; cut regression time by ~70%.
  • Validated LLM outputs with 200+ DeepEval and Pytest cases; reduced false failures from ~15% to ~5%.
insurance · rpaDec 2022 – Mar 2025

Senior QA Analyst – Team Lead

Capgemini · Farmers Insurance, Guidewire PolicyCenter
  • Led end-to-end QA for P&C insurance applications and mentored the QA team.
  • Designed and ran UiPath RPA test automation with Zephyr and Jira traceability.
salesforce · clmDec 2019 – Dec 2022

Senior QA Analyst

IBM · Apttus CLM (Salesforce) & Compliance Centre
  • Owned functional, regression, integration and UAT testing; led defect triage in Jira.
  • Reduced post-production defects through early gap analysis and walkthroughs.
fintech · initSep 2015 – Apr 2019

QA Analyst

Bhanix Financial · CASHe
  • Web, API (Postman, Swagger), SQL data and mobile testing for a FinTech lending app.
06 / credentials

Certifications & education

AI Agents Fundamentals

Hugging Face · Apr 2026

UiPath Certified Automation Developer Associate

UiPath

SAFe 6.0 Agile Certified

Scaled Agile

Analyzing and Visualizing Data with Power BI

Microsoft

B.Tech, Electrical & Electronics Engg.

JNTU Hyderabad · 2007–2011

$ ./connect --role "AI QA"

Let's talk about AI quality

I'm looking for roles in AI testing, LLM evaluation and test automation. Reach me by email, phone or LinkedIn.