Senior Automation QA; JS) Engineer
Listed on 2026-10-09
-
Software Development
AI QA / Validation Engineer, Software Testing
Ciklum is looking for a Senior Automation QA (JavaScript) Engineer to join our team full-time in India.
We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.
About the role:As a Senior Automation QA (JavaScript) Engineer, become a part of a cross-functional development team engineering experiences of tomorrow. Ciklum is seeking a Senior AI Automation QA Engineer to join the client project, an ambitious initiative to build an enterprise-grade Agent Development Platform (ADP) that our client will commercialize to their own customers. This is quality engineering for a system that is non-deterministic by nature, where "expected output" is a statistical claim rather than a string match, and where the platform's reliability guarantees no lost workflow runs, strict tenant isolation, graceful degradation under provider outages are contractual commitments, not aspirations.
You will build the automation that proves those guarantees hold, release after release.
You will own the test automation that spans the whole platform: the Frontend and Backend that agent developers depend on, the long-running Temporal workflows that must survive worker crashes and provider failures, the multi-tenant isolation boundaries that protect one customer's data from another, and the performance and concurrency targets defined in the platform's NFRs. Much of your craft is adversarial and failure-first injecting worker kills, retry storms, and malformed documents to verify the system recovers exactly as specified, and negative-testing disallowed tool calls and egress paths to confirm the sandbox holds.
Where agent behavior must be judged rather than asserted, you will partner closely with the evaluation function, contributing the test infrastructure, synthetic-data pipelines, and CI integration that make quality measurable and repeatable.
This is a high-visibility role where we establish the production-ready foundation the rest of the program is built on. Your automated suites become release gates, and the evidence they produce feeds directly into client milestone sign-offs so your work must be reproducible, trustworthy, and defensible to a sophisticated client. As the platform scales toward an "agent-a-month" delivery cadence, you will build the reusable test scaffolds and harnesses that let quality keep pace with delivery, ensuring that month twelve ships as safely as month one.
You will work shoulder-to-shoulder with AI engineers, platform engineers, and the evaluation team, embedding quality from the design stage rather than inspecting for it at the end.
- Design, build, and maintain automated test suites for the platform, Frontend and Backend, including contract tests that protect the agent-developer experience as the SDK evolves
- Build automated tests for long-running Temporal workflows including failure injection (worker kill, provider outage, retry storms) to verify durable-execution guarantees such as zero lost runs and correct compensation/saga behavior
- Automate multi-tenant isolation verification cross-tenant data access, configuration bleed, and cost-attribution correctness executed per release as a platform acceptance criterion
- Verify agent execution sandboxing and egress controls through negative testing of disallowed tool calls and network destinations
- AI-Specific & Non-deterministic Testing
- Design test strategies for LLM-driven behavior: statistical assertions over repeated runs (pass^k consistency), semantic-similarity scoring, confidence-threshold validation, and flakiness quarantine that distinguishes model variance from genuine regression
- Build and continuously extend adversarial test corpora document-borne prompt injection in lease PDFs, tool-call hijack, and data-exfiltration probes aligned to OWASP LLM Top 10 and recognized red-team taxonomies
- Verify Human-in-the-Loop (HITL) escalation behavior confidence thresholds fire correctly and no gated/irreversible action completes without authorization
- Partner with the evaluation function on golden-dataset-driven checks, contributing test infrastructure, synthetic-data (synthetic lease) pipelines, and CI integration that make agent quality measurable and repeatable
- Performance, Reliability & NFR…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).