Demo

MTS - Kernel Engineer

Acceler8 Talent
Mountain View, CA Full Time
POSTED ON 8/2/2026
AVAILABLE BEFORE 8/30/2026

Kernel Engineer — Compute and Accelerators


Mountain View, CA


About the Role


An early-stage AI hardware company is seeking a Kernel Engineer to develop and optimize specialized compute kernels for a custom accelerator platform.


This role sits at the critical boundary between machine learning workloads and silicon. Your work will directly influence how efficiently the hardware executes tensor operations, moves data, and uses the underlying memory hierarchy.


You will work closely with architecture, compiler, simulation, and systems teams to define the kernel programming model, implement core operations, and build the profiling workflows used to evaluate hardware and software performance.


What You’ll Do


  • Develop and optimize compute kernels for a custom AI accelerator.
  • Implement tensor operations, data-movement patterns, and memory-hierarchy optimizations.
  • Build and maintain profiling infrastructure to measure kernel performance against architectural targets.
  • Define execution and data-shuffling patterns across general-purpose control cores, tensor-processing units, and specialized compute engines.
  • Help shape the kernel programming model, including thread execution, register-passing conventions, synchronization, and memory-management strategies.
  • Enable end-to-end kernel execution in simulation and pre-silicon environments.
  • Collaborate with compiler engineers on intermediate representations, lowering strategies, and kernel integration.
  • Use kernels as validation targets for compiler and architecture development.
  • Create technical documentation, examples, and kernel-development guides for the broader engineering team.
  • Investigate performance bottlenecks and recommend hardware or software improvements.


What We’re Looking For


  • Strong C and C programming skills with experience writing production-quality systems or performance-critical code.
  • Deep experience with CUDA or a comparable accelerator-programming model.
  • Strong understanding of:
  • Warp, wavefront, or thread-group execution
  • Memory coalescing
  • Shared or local memory
  • Registers and caches
  • Synchronization
  • Data locality
  • Memory bandwidth and latency
  • Ability to reason about computer architecture, including pipelines, execution units, memory hierarchies, and data-movement costs.
  • Strong performance-profiling and optimization experience.
  • Experience identifying bottlenecks, measuring throughput and latency, and iterating until performance targets are met.
  • Practical understanding of tensor and numerical operations, including:
  • GEMM
  • Convolution
  • Attention
  • Reductions
  • Scatter and gather
  • Elementwise operations
  • Python experience for scripting, tooling, automation, and integration work.
  • Ability to collaborate effectively across architecture, compiler, and hardware teams.


Preferred Experience


  • Triton, CUTLASS, or similar kernel-development frameworks.
  • MLIR, LLVM, or compiler infrastructure.
  • RISC-V, x86, ARM64, or another instruction-set architecture.
  • High-performance computing or scientific computing.
  • Custom ASIC, GPU, NPU, or accelerator software.
  • FPGA development or experience reading RTL.
  • Verilog or SystemVerilog.
  • Architectural simulators or instruction-set simulators.
  • Kernel DSL design.
  • Hardware-software co-design.


Relevant Keywords

Kernel Engineer, Compute Kernels, Accelerator Kernels, GPU Kernels, CUDA, CUDA C , C , Python, Triton, CUTLASS, Tensor Operations, GEMM, Matrix Multiplication, Convolution, Attention, Reductions, Scatter/Gather, Elementwise Operations, Parallel Programming, SIMT, SIMD, Warp Execution, Wavefront Execution, Thread Blocks, Memory Coalescing, Shared Memory, Registers, Cache Optimization, Memory Hierarchy, Data Locality, Data Movement, Synchronization, Performance Profiling, Performance Optimization, Throughput, Latency, Roofline Analysis, Nsight Compute, Nsight Systems, Computer Architecture, Custom ASIC, AI Accelerator, NPU, GPU, MLIR, LLVM, Kernel DSL, Compiler Integration, Architectural Simulation, Instruction Set Simulator, RISC-V, ARM64, x86, HPC, Scientific Computing, Verilog, SystemVerilog, FPGA, Hardware-Software Co-Design, Member of Technical Staff, MTS, PMTS, Principal Member of Technical Staff


Salary : $260,000 - $320,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a MTS - Kernel Engineer?

Sign up to receive alerts about other jobs on the MTS - Kernel Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$85,996 - $102,718
Income Estimation: 
$111,859 - $131,446
Income Estimation: 
$110,457 - $133,106
Income Estimation: 
$105,809 - $128,724
Income Estimation: 
$122,763 - $145,698
Income Estimation: 
$93,348 - $109,523
Income Estimation: 
$112,230 - $133,397
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Acceler8 Talent

  • Acceler8 Talent Boston, MA
  • Entry Level Recruiter - 2025 Grads - 2026 Start! Do you want to join an organization with a 5-star Glass Door review? Did you graduate in 2025? Interested ... more
  • 11 Days Ago

  • Acceler8 Talent Mountain View, CA
  • Compiler Engineer We are seeking a Compiler Engineer to join a scale up team founded by ex-Tesla AI leaders, taking a lead in designing and building their ... more
  • 11 Days Ago

  • Acceler8 Talent Berkeley, CA
  • Senior Digital IC Engineer Berkeley, CA (Onsite) Up to 250k base founding equity. I'm working with an early-stage semiconductor startup building a fundamen... more
  • 11 Days Ago

  • Acceler8 Talent San Francisco, CA
  • ML Researcher (World Models) We are seeking Researchers with expertize in World Models to join an an NVIDIA Inception & SPC backed lab building SOTA AI voi... more
  • 11 Days Ago


Not the job you're looking for? Here are some other MTS - Kernel Engineer jobs in the Mountain View, CA area that may be a better fit.

  • NIO San Jose, CA
  • JOB DESCRIPTION About NIO NIO is a pioneer and a leading company in the premium smart electric vehicle market. Founded in November 2014, NIO's mission is t... more
  • 12 Days Ago

  • Black Sesame Technologies Inc San Jose, CA
  • Role We are looking for a Senior NPU Kernel/Operator Engineer to lead the design and optimization of high-performance kernels for a custom AI accelerator /... more
  • 23 Days Ago

AI Assistant is available now!

Feel free to start your new journey!