Location: Remote (LATAM)
Rate: $1,500 - $2,000/month (Half-Time)
Company: One of Our Amazing Partners
ABOUT THE COMPANY
One of Our Amazing Partners is a SaaS company building AI-powered experiences for the B2B services industry. Their platform uses relationship intelligence, signal detection, and agentic workflows to help companies grow smarter. The engineering team is actively shipping new product and the quality systems built now will define how the platform scales.
WHAT WE'RE LOOKING FOR
A product-minded engineer who can turn business intent into executable evidence. You will define what good behavior means across both deterministic software and AI-powered agent outputs, create realistic personas and scenarios, build automation, and evaluate results where exact-text assertions are not sufficient. This is not downstream manual testing. It is a hands-on engineering role that reports into the CTO organization and works as an embedded partner to Product, Design, and Engineering from discovery through release and production learning.
IN THIS ROLE, YOU WILL
- Translate product decisions, business rules, and agent specifications into clear behavioral contracts and acceptance criteria
- Create representative firm and operator personas, together with their contacts, relationship networks, target organizations, integrations, consent states, and data conditions
- Build and maintain critical-path automation across browser, API, worker, integration, and data layers, including Playwright-based journeys
- Design AI evaluation suites that combine deterministic checks, structured rubrics, repeated trials, calibrated graders, and periodic human review
- Test signal detection, relationship-sensitive routing, scoring, recommendations, grounding, memory, permissions, and graceful failure across realistic and adversarial conditions
- Define release evidence and production quality signals so that go/no-go decisions are driven by measurable evidence rather than isolated demos or intuition
- Turn operator corrections and escaped failures into durable regression cases that permanently strengthen the system
- Improve the shared simulation, scenario, trace, and evaluation tooling used by Product and Engineering
- Ensure failures are diagnosable across ingestion, detection, routing, agent decision, persistence, and presentation
YOU'LL SUCCEED IF YOU HAVE
- Strong experience in software quality, test engineering, product engineering, developer productivity, or AI evaluation for complex production systems
- Hands-on experience evaluating LLM applications or agentic systems, including tool use, retrieval, grounding, memory, and non-deterministic outputs
- Strong TypeScript and/or Python skills with experience building maintainable test or evaluation infrastructure
- Practical depth in Playwright or equivalent browser automation, plus API, integration, contract, worker, and data-pipeline testing
- Ability to design gold datasets, rubrics, sampling strategies, automated graders, and human calibration workflows
- Strong product and business judgment: you can assess whether an output is useful to a real user, not only whether it is syntactically valid
- Comfort diagnosing asynchronous and distributed systems with retries, partial failure, eventual consistency, and multiple data stores
- Clear cross-functional communication and the ability to create alignment in ambiguous, fast-moving product work
NICE TO HAVE
- B2B SaaS, sales or relationship intelligence, professional services, graph analytics, recommendation systems, or calibrated scoring
- Experience establishing a quality or evaluation discipline at an early-stage company and mentoring others in quality practices
- Model and prompt versioning, offline and online evaluation, observability, and cost-quality tradeoff experience
- Familiarity with buyer-intent signal detection or relationship-sensitive data systems
WHY JOIN THIS TEAM
You will be the person who defines what quality means for an AI product in its most formative stage. The team is small enough that your decisions shape the product directly, and the problem space (relationship intelligence, agentic workflows, real-time signal detection) is genuinely interesting engineering territory. If you want to build the evaluation and quality systems for an AI-native platform from the ground up, this is the role.
ABOUT TALENT SCOUT
Talent Scout connects top creative, marketing, and operations professionals across Latin America with leading U.S. brands and agencies. We specialize in long-term, full-time placements, meaning that when you join through Talent Scout, you’re not stepping into a freelance or temporary role — you’re becoming a valued member of a client’s team, with support every step of the way.
To learn more about working with Talent Scout, please visit our Careers Page and follow us on LinkedIn and Instagram.