AI Testing Specialist
Job in
New York City, Richmond County, New York, USA
Listed on 2026-07-29
Listing for:
Ova Technologies
Full Time
position Listed on 2026-07-29
Job specializations:
-
Software Development
AI QA / Validation Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Job Title
AI Testing Specialist
LocationHybrid / Remote
Employment TypeFull-time
Job SummaryWe are seeking an AI Testing Specialist to ensure the quality, reliability, security, and performance of AI-powered applications, machine learning models, and generative AI solutions. The ideal candidate will develop and execute comprehensive testing strategies for AI systems, validate model outputs, assess AI-specific risks, and collaborate with cross-functional teams to deliver high-quality AI products.
Key Responsibilities- Design and execute test strategies for AI, machine learning, and generative AI applications.
- Create and maintain test plans, test cases, and test data for AI features and workflows.
- Validate AI model outputs for accuracy, consistency, relevance, factuality, and reliability.
- Evaluate AI systems for hallucinations, bias, toxicity, fairness, and robustness.
- Perform functional, regression, integration, API, end-to-end, performance, usability, and security testing.
- Test prompt-based applications and optimize prompts for consistent results.
- Develop automated testing frameworks for AI applications and APIs.
- Verify data quality, preprocessing pipelines, and model inputs.
- Conduct stress, load, and scalability testing for AI services.
- Identify, document, prioritize, and track defects using bug management tools.
- Collaborate with AI engineers, data scientists, software developers, product managers, and UX teams.
- Monitor production AI systems and support continuous quality improvement.
- Prepare test reports, quality metrics, and release recommendations.
- Ensure compliance with organizational AI governance, security, privacy, and regulatory requirements.
- Bachelor's degree in Computer Science, Information Technology, Software Engineering, Data Science, or a related field.
- 3–5+ years of experience in software testing, QA, or AI testing.
- Strong understanding of software testing methodologies and quality assurance principles.
- Experience testing APIs, web applications, and cloud-based systems.
- Familiarity with AI, machine learning, and generative AI concepts.
- Experience working in Agile or Scrum environments.
- Experience testing Large Language Model (LLM) applications.
- Knowledge of prompt engineering and AI evaluation methodologies.
- Experience with Responsible AI practices and AI governance.
- AI, cloud, or software testing certifications.
- Experience with MLOps workflows and model lifecycle management.
- Manual and automated testing
- Test planning and execution
- API testing (Postman, REST Assured)
- Automation frameworks (Selenium, Playwright, Cypress)
- Programming (Python, Java, JavaScript, or C#)
- SQL and database validation
- Git and CI/CD tools
- Test management tools (Jira, Test Rail, Zephyr)
- Performance testing (JMeter, k6, Load Runner)
- AI model evaluation techniques
- Prompt engineering
- LLM testing and validation
- AI safety testing (hallucinations, bias, toxicity, prompt injection, jailbreak resistance)
- Data validation and preprocessing verification
- JSON, REST APIs, and cloud platforms (AWS, Azure, Google Cloud)
- Analytical and critical thinking
- Strong attention to detail
- Problem-solving
- Effective communication
- Collaboration across multidisciplinary teams
- Documentation and reporting
- Adaptability
- Time management
- Continuous learning
- AI-powered enterprise applications
- Conversational AI and chatbots
- Generative AI products
- Machine learning platforms
- Retrieval-Augmented Generation (RAG) systems
- SaaS and cloud-native applications
- Healthcare, finance, retail, or other regulated industries
- Test coverage and automation coverage
- Defect detection and prevention rate
- AI response quality and reliability
- Reduction in production defects
- Model evaluation accuracy
- Compliance with AI quality and governance standards
- Release readiness and stability
- Customer satisfaction and user experience
- Test execution efficiency
- AI evaluation frameworks (Deep Eval, Ragas, Lang Smith, Promptfoo)
- Lang Chain or similar AI orchestration frameworks
- Vector databases
- Docker and Kubernetes
- MLOps tools (MLflow, Kubeflow, Sage Maker)
- Explainable AI (XAI) concepts
- Data annotation and synthetic data generation
- Accessibility testing
- Security testing for AI systems
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×