Posted at: 22 July
Senior Performance Engineer
Company
NVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.
Remote Hiring Policy:
NVIDIA supports flexible remote work arrangements and hires from various regions globally, including the Americas, Europe, Asia, and the Middle East, with roles that may require collaboration across time zones.
Job Type
Full-time
Allowed Applicant Locations
Israel
Job Description
NVIDIA is seeking a highly skilled Senior Performance Engineer to join our Performance and R&D organizations. In this role, you will help build and evolve systems that support performance analysis, telemetry, and optimization for large-scale GPU- and CPU-based clusters used in AI and high-performance computing environments. You will work closely with hardware, networking, firmware, and software teams to collect, analyze, and interpret performance data from live systems. This is a fast-paced R&D environment where system behavior and requirements evolve rapidly, requiring adaptable engineering solutions and strong analytical thinking.What you’ll be doing:Profile, benchmark, and analyze AI and HPC workloads on GPU and CPU clustersExplore performance characteristics of high-performance networking and collective communications (e.g., NCCL, RDMA, MPI, RoCE)Identify performance bottlenecks across networking, compute, memory, and system architectureDevelop and enhance performance analysis, benchmarking, and diagnostic toolsDefine performance test plans and establish expectations for new technologies and platformsCollaborate across hardware, firmware, networking, systems, and software teams to provide actionable performance insightsSupport telemetry collection and data refinement efforts to enable accurate performance analysisMaintain high standards for data quality, reproducibility, and traceability of performance resultsWhat we need to see:B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent experience5+ years of experience in performance analysis, systems engineering, or HPC/AI infrastructureDemonstrated expertise in performance analysis skills and methodologiesHands-on experience with high-performance networking (RDMA, MPI, NCCL, congestion control)Strong understanding of system performance metrics (latency, throughput, resource utilization)Exposure to hardware, firmware, or embedded telemetry environmentsStrong analytical, problem-solving, and communication skillsAbility to work effectively in cross-functional, fast-paced R&D teamsWays to stand out from the crowd:Knowledge of CUDA, NCCL internals, and congestion control algorithmsDeep system-level understanding of CPU architectures, GPUs, HCAs, memory, and PCIeExperience with NVIDIA GPUs, CUDA, and deep learning frameworks such as PyTorch or TensorFlowExperience with cloud platforms Proficiency in Python; experience with Bash and C/C++ is a plus as well as a strong experience working in Linux environments