QA Lead
HugeInc
| Company | HugeInc |
| Category | Engineering |
| Location | Colombia |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Lead |
| Salary | Not stated by the employer |
| Posted | 2 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
Location: This position is remote within Colombia
About the Role
We are seeking a seasoned, battle-tested QA Lead & AI Brand Evaluation Strategist to take full ownership of the quality, integrity, and strategic evaluation of our rapidly expanding suite of agentic and AI-powered product features.
This is not a traditional, checklist-driven QA role. You will operate in a highly dynamic, non-deterministic environment, leading the strategy to evaluate probabilistic software at scale. Your core mission is to transform complex, ambiguous business requirements into robust evaluation criteria and scoring systems. You will lead the design of tools and frameworks that process massive volumes of data, ensuring our brand scoring metrics are precise, reliable, and actionable.
If you have a proven track record of leading QA initiatives, navigating technical ambiguity with leadership, and building evaluation architectures for complex data ecosystems, this role is for you.
What you'll do
Strategic QA Leadership: Own the end-to-end testing and evaluation strategy for multi-agent architectures and high-volume data pipelines, defining direction even when product requirements are fluid or ambiguous.
Business-to-Evaluation Translation: Act as the primary bridge between Product, Data Science, and Engineering. Translate complex business objectives into concrete, data-driven evaluation criteria and operational rubrics.
Brand Scoring Architecture: Oversee the methodology for scoring brands against complex criteria, ensuring that automated systems accurately transform massive datasets into precise, reliable business metrics (e.g., Share of Model, Net Sentiment).
Framework & Tool Building: Drive the creation of internal tools, data-driven testing pipelines, and validation frameworks capable of handling non-deterministic AI outputs and large-scale information ingestion.
System Integrity & Governance: Establish risk-adaptive guardrails and checkpoints to ensure data precision, compliance (PII), and logical reasoning across complex, multi-step agentic workflows.
Lead & Orchestrate Evaluation Frameworks: Define, architect, and oversee validation processes for application layers, complex API workflows, and the data transformation engine.
Define Brand Scoring Criteria: Establish the mathematical and logical rules that operate over massive data volumes to audit whether brand evaluations match underlying raw metrics reliably.
Direct "Golden" Test Set Curation: Lead the strategy for compiling, synthesizing, and maintaining massive baseline validation datasets to test edge cases, user intent, and model drift.
Drive Adversarial & Ambiguity Testing: Design "red-teaming" scenarios and boundary testing to evaluate how gracefully the system handles prompt variations, conversational drift, and highly ambiguous data inputs.
Tooling Innovation: Collaborate with engineering to build proprietary testing tools or automation scripts that streamline the ingestion and analysis of high-volume data for QA purposes.
Defect Profiling & Root-Cause Strategy: Move beyond simple bug logging; analyze trends in statistical behavioral defects and guide cross-functional teams toward systemic root-cause resolution.
Platform Observability & Metrics: Monitor cloud logging and data observability dashboards to track execution latency, data drift, and accuracy trends, transforming insights into platform optimizations.
What we're looking for
Experience: 8+ years of proven experience in QA Engineering, Systems Analysis, or Software Testing, with at least 2+ years in a Lead or Strategic role managing complex, data-heavy digital platforms.
Leadership in Ambiguity: Proven "cancha" (battle-tested experience) leading QA initiatives in fast-paced environments where requirements are not fully defined, showing a high capacity to resolve ambigui