Head of QA
IO Tech Solutions Hong KongHead of QA
You will lead Quality & Reliability Engineering (QRE) — the team responsible for product quality
and production reliability across SleekFlow. You own the testing strategy, the test platform,
release quality gates, observability and diagnostic tooling, incident management, and availability
reporting. You report to engineering leadership.
What you'll own
1. Testing Strategy & Platform
● Framework Ownership: Builds and maintains the E2E frameworks, CI integration, and
test-data management used daily by application engineers.
● Automated & Proactive Testing: Drives continuous, multi-tiered regression testing based
on real production patterns, and leads pre-release and exploratory testing programs.
● Shared Responsibility: The platform team provides the infrastructure and framework,
while application engineers write the actual feature tests.
2. Release Quality & Delivery
● Policy Setting: Defines the release-quality policy (which gates a release, which test
suites run, and how rollbacks are triggered) based on the change's risk level.
● Pipeline Enforcement: Owns what the pipeline enforces, while release engineers
manage pipeline execution.
3. Observability & Diagnostic Tooling
● Platform & Standards: Owns the infrastructure and standards for structured logging,
metrics, tracing, and alerting to ensure production issues are easily diagnosable.
● Data Pipelines: Manages the tools and pipelines, while separate application and platform
teams handle the actual instrumentation of their services.
4. Incident Management
● Lifecycle Ownership: Coordinates the entire response for all production incidents
(severity classification, stakeholder communication, and postmortems) regardless of the
root cause.
● Continuous Learning: Tracks post-incident follow-ups and feeds lessons learned back
into testing and monitoring improvements. (The team that caused the issue owns the fix).
5. Availability, Stability, & Reliability
● Target Definition: Co-defines system availability, degradation, and downtime targets with
leadership.
● Metrics & Reporting: Translates definitions into measurable SLIs and health signals,
builds operational dashboards, and produces regular reliability reports to guide quality
investments.
What we’re looking for
10+ years in software engineering, quality engineering, or production reliability, with at least 3–4
years in a leadership role (managing people, setting org-level strategy, influencing cross-team
decisions).
You’ve worked across both quality engineering and production reliability — either in separate
roles or combined. You understand how product quality and site reliability reinforce each other
and can explain that to an engineering org.
Specific experience we value:
● Designing tiered regression strategies — deciding what to run per repo or change class
in CI/CD so delivery stays fast without blind spots.
● Running incident response end-to-end — from detection to postmortem to tracked
prevention — for a multi-service, multi-region SaaS product.
● Building or scaling a test automation platform used by other engineers daily (E2E
frameworks, CI pipelines, test-data management, flake policy).
● Building or owning observability infrastructure — structured logging, distributed tracing,
metrics pipelines, and alerting — that engineers rely on daily to diagnose production
issues.
● Defining availability or reliability metrics (SLIs, SLOs, or equivalent) and making them
mean something to leadership — not just dashboards that exist but that nobody trusts.
● Operating in a fast-shipping engineering culture where features are built in days, specs
are often thin, and the quality model has to keep pace.
● You translate between “5xx on service X” and “inbox is broken for customers.” You
present quality and availability to leadership in business terms and work with engineers
in technical terms. You’re comfortable being the person who says “this isn’t ready to
ship” when it matters.
Nice to have
● Hands-on with Playwright or similar modern E2E frameworks and CI-driven test
pipelines.
● Experience with multi-region SaaS (Azure preferred, but cloud-native fundamentals
transfer).
● Experience with structured pre-release testing programs (internal beta, dogfooding) at a
product company.
● Experience co-defining customer-facing SLAs or availability reporting with product or
leadership.