Sr. AI Validation Engineer

Location: Portland, Oregon
Category: Software Development
Employment Type: Contract
Work Location: Remote
Job ID: 36299
Date Added: 08/31/2026

We are seeking a Senior AI Infrastructure Validation Engineer to validate the readiness, reliability, and performance of large-scale AI infrastructure environments before they reach production.

This role sits at the intersection of systems engineering, networking, GPU infrastructure, container platforms, and distributed AI workloads. The focus is not simply executing predefined test cases, but designing validation strategies that uncover failure modes early, accelerate root-cause isolation, and provide clear evidence that clusters are ready for production workloads.

The ideal candidate approaches infrastructure validation as an engineering discipline, with the ability to understand how compute, networking, storage, orchestration, firmware, and distributed workloads interact across complex AI environments.

Responsibilities

  • Design and execute comprehensive validation plans for AI infrastructure spanning:
    • Compute nodes
    • GPU communication
    • Network fabric health
    • Storage access
    • Container orchestration
    • Distributed workload readiness
  • Perform structured bring-up, soak, regression, qualification, and certification testing for new and modified AI cluster environments.
  • Reproduce, troubleshoot, and isolate failures involving:
    • Distributed training workloads
    • GPU/node instability
    • NCCL and communication libraries
    • Kubernetes and container platforms
    • Storage paths and throughput
    • Ethernet and InfiniBand transport behavior
  • Validate Ethernet and InfiniBand environments, including host readiness, RDMA connectivity, and end-to-end workload behavior.
  • Correlate failures across system logs, telemetry, firmware state, infrastructure health, and workload symptoms to accelerate root-cause analysis.
  • Determine whether failures originate from infrastructure, hardware, networking, software frameworks, workloads, or configuration.
  • Partner with Linux, networking, deployment, platform, and infrastructure engineering teams to close validation gaps before production handoff.
  • Define defect signatures, validation criteria, pass/fail thresholds, readiness assessments, and release recommendations.
  • Build and improve automation for:
    • Cluster certification
    • Health scoring
    • Regression testing
    • Evidence collection
    • Post-change validation
    • Readiness reporting
  • Improve repeatability and scalability of validation processes across large AI infrastructure environments.

Required Qualifications

  • 7+ years of experience in infrastructure validation, systems validation, performance engineering, infrastructure QA, HPC, or AI infrastructure certification.
  • Strong hands-on troubleshooting experience across:
    • Linux systems
    • GPU infrastructure
    • Network fabrics
    • Containers
    • Distributed workloads
  • Demonstrated experience designing validation strategies, frameworks, and qualification approaches, rather than only executing predefined test cases.
  • Understanding of AI infrastructure dependencies including:
    • NCCL
    • RDMA
    • Storage throughput
    • Multi-node communication
    • Distributed training behavior
    • Cluster orchestration
  • Ability to distinguish infrastructure defects from application, framework, workload, and configuration issues.
  • Strong scripting and automation skills using Python, Bash, or similar technologies.
  • Experience automating test execution, evidence collection, analysis, and reporting.
  • Strong root-cause analysis and systems troubleshooting skills.
  • Excellent written communication skills with experience producing detailed defect reports, readiness assessments, and technical recommendations.

Preferred Qualifications

  • Experience validating GPU clusters, large-scale AI training environments, HPC systems, or AI infrastructure prior to production deployment.
  • Experience with burn-in, soak testing, telemetry analysis, and cluster qualification.
  • Familiarity with hardware, firmware, driver, and software compatibility testing.
  • Experience building validation suites that support both:
    • Pre-production deployment readiness
    • Ongoing production health and regression validation
  • Experience supporting high-density GPU or distributed computing environments.

Tools & Technologies

Infrastructure & Platforms

  • Linux
  • GPU Infrastructure
  • Kubernetes
  • Containers
  • Distributed Training Systems

Networking

  • Ethernet
  • InfiniBand
  • RDMA
  • NCCL

Automation

  • Python
  • Bash
  • Automation Frameworks
  • Cluster Certification Frameworks

Validation & Observability

  • Telemetry Platforms
  • Performance Analysis Tools
  • Infrastructure Validation Tools
  • Burn-In / Soak Testing
  • Health Monitoring
  • Regression & Qualification Testing

Apply Now

Fill out the form below to submit your information for this opportunity. Please upload your resume as a doc, pdf, rtf or txt file. Your information will be processed as soon as possible.

Please upload your resume as a doc, pdf, rtf or txt file.

Related Jobs