GPU Kernel Expert with 3+ years of experience in developing and optimizing GPU kernels, focusing on AI systems performance and numerical correctness.
Compensation
$146K–$187K
yearly · USD
Experience
3–15 yrs
Location
Remote
United States
Compensation
$146K–$187K
yearly · USD
Experience
3–15 yrs
Location
Remote
United States
The Brief
TITLE
GPU Kernel Expert
TYPE
Contract
POSTED
Aug 27, 2026
JOB ID
01a04257
TITLE
GPU Kernel Expert
TYPE
Contract
POSTED
Aug 27, 2026
JOB ID
01a04257
Remote | Independent Contractor | United States | $145,600–$187,200 annualized ($70–$90/hour)
Apply your GPU kernel development and optimization expertise to help improve the capabilities of advanced AI systems.
As a GPU Kernel Expert, you’ll evaluate GPU and accelerator kernel development tasks used to train and assess frontier AI models. You’ll assess numerical correctness, performance benchmarking, task scope, compilation, and runtime behavior across a range of kernel development scenarios.
Your technical judgment will help ensure that AI training and evaluation tasks accurately reflect real-world GPU and accelerator engineering challenges.
Evaluate GPU and accelerator kernel tasks for quality, correctness, completeness, and technical rigor.
Assess numerical correctness using appropriate absolute, relative, and ULP tolerance criteria.
Evaluate whether reference implementations are appropriate and reliable for validating kernel outputs.
Review performance benchmarks for fairness, consistency, and methodological soundness.
Assess task scope and determine whether requirements are technically clear, realistic, and appropriately defined.
Evaluate compilation and runtime behavior across different hardware and software environments.
Identify common failure modes, including:
Driver and environment mismatches
Out-of-memory (OOM) conditions
Kernel launch configuration errors
Shape and stride mismatches
Autotuning failures
Provide clear, structured, rubric-based written feedback.
Evaluate kernel tasks involving generation, translation, migration, debugging, optimization, and operator fusion.
3+ years of hands-on experience developing, optimizing, or verifying GPU or accelerator kernels.
Professional experience with at least two of the following:
CUDA
Triton
NKI
Pallas (JAX)
Strong understanding of numerical correctness for GPU kernels, including:
Absolute and relative tolerances
ULP-based comparisons
Reference implementation selection
Demonstrated experience with performance profiling and benchmarking using tools or methodologies such as:
NVIDIA Nsight
Nsight Compute (NCU)
Roofline analysis
Framework-native profilers
Familiarity with common GPU/kernel compilation and runtime failure modes.
Experience with at least three of the following kernel task types:
Generation from specification
Translation or lowering across frameworks
Migration between hardware targets
Kernel debugging
Performance optimization
Operator fusion
Strong analytical skills and attention to technical detail.
Ability to communicate complex kernel engineering concepts clearly in writing.
Experience in the following areas is highly valuable:
Experience across both NVIDIA GPU ecosystems such as CUDA/Triton and custom accelerator ecosystems such as NKI/Pallas/TPU.
Background in compiler engineering, MLIR, or intermediate-representation lowering.
Deep understanding of memory-hierarchy optimization, including:
Shared-memory tiling
Register pressure
Bank conflicts
Memory coalescing
Contributions to GPU or accelerator kernel libraries and related open-source projects.
Experience with technologies such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls.
Rate: $70–$90/hour
Annualized Equivalent: $145,600–$187,200
Location: United States
Work Arrangement: Fully remote
Engagement Type: Independent contractor
Schedule: Flexible
Payment: Weekly via Stripe or Wise
Annualized compensation is based on 2,080 hours per year for comparison purposes only. Actual earnings depend on the number of hours and projects completed.
Apply your GPU and accelerator engineering expertise to advanced AI development.
Work on technically challenging kernel evaluation and optimization problems.
Help improve how AI systems understand low-level performance, numerical correctness, and hardware acceleration.
Evaluate realistic engineering scenarios across GPUs, accelerators, frameworks, and compilers.
Work remotely with a flexible schedule.
Contribute specialized expertise to high-impact AI training and evaluation projects.
You will be engaged as an independent contractor.
Work is fully remote and can be completed on your own schedule.
Projects may be extended, shortened, or concluded early depending on project needs and performance.
Your work will not require access to confidential or proprietary information belonging to any employer, client, or institution.
Payments are made weekly via Stripe or Wise based on services rendered.
H-1B and STEM OPT candidates cannot be supported at this time.
All qualified applicants will be considered without regard to legally protected characteristics. Reasonable accommodations are available upon request.
About the company
Recruitment Room is a global workforce solutions company helping businesses build, manage, and scale distributed teams across international markets. We partner with startups, scale-ups, SMEs, and enterprise organizations to deliver end-to-end workforce solutions that support business growth, operational efficiency, and global expansion.
Our services extend beyond talent acquisition to include Employer of Record (EOR), Contractor of Record (COR), contractor management, global payroll, HRIS, workforce strategy, recruitment process outsourcing (RPO), executive search, and AI-powered talent intelligence. By combining human expertise with intelligent technology, we simplify the complexities of hiring, employing, and managing talent across borders.
For professionals, Recruitment Room provides access to career opportunities with innovative employers worldwide while supporting long-term career growth through our global talent ecosystem.