{"absolute_url":"https://job-boards.greenhouse.io/xai/jobs/4533894007","data_compliance":[{"type":"gdpr","requires_consent":false,"requires_processing_consent":false,"requires_retention_consent":false,"retention_period":null,"demographic_data_consent_applies":false}],"internal_job_id":4327690007,"location":{"name":"Palo Alto, California"},"metadata":[{"id":22659078007,"name":"Job Description via Google Doc URL [please share with Olivia]","value":null,"value_type":"short_text"},{"id":16340689007,"name":"Featured Role","value":null,"value_type":"yes_no"}],"id":4533894007,"updated_at":"2026-08-27T18:09:48-04:00","requisition_id":"75","title":"Software Engineer - Training/Inference (C++)","company_name":"SpaceXAI","first_published":"2024-10-04T20:14:13-04:00","language":"en","application_deadline":null,"content":"\u0026lt;div class=\u0026quot;content-intro\u0026quot;\u0026gt;\u0026lt;p\u0026gt;\u0026lt;span style=\u0026quot;font-family: arial, helvetica, sans-serif;\u0026quot;\u0026gt;SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.\u0026amp;nbsp;\u0026lt;/span\u0026gt;\u0026lt;span style=\u0026quot;font-family: arial, helvetica, sans-serif;\u0026quot;\u0026gt;Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. \u0026lt;/span\u0026gt;\u0026lt;span style=\u0026quot;font-family: arial, helvetica, sans-serif;\u0026quot;\u0026gt;We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. \u0026lt;/span\u0026gt;\u0026lt;span style=\u0026quot;font-family: arial, helvetica, sans-serif;\u0026quot;\u0026gt;All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.\u0026lt;/span\u0026gt;\u0026lt;/p\u0026gt;\u0026lt;/div\u0026gt;\u0026lt;h3\u0026gt;About the Role:\u0026lt;/h3\u0026gt;\n\u0026lt;ul\u0026gt;\n\u0026lt;li\u0026gt;We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency).\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale\u0026lt;/li\u0026gt;\n\u0026lt;/ul\u0026gt;\n\u0026lt;p\u0026gt;\u0026lt;span style=\u0026quot;font-size: 14pt;\u0026quot;\u0026gt;\u0026lt;strong\u0026gt;Responsibilities:\u0026amp;nbsp;\u0026lt;/strong\u0026gt;\u0026lt;/span\u0026gt;\u0026lt;/p\u0026gt;\n\u0026lt;ul\u0026gt;\n\u0026lt;li\u0026gt;Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Optimize latency and throughput of model inference under real production workloads.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation).\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.\u0026lt;/li\u0026gt;\n\u0026lt;/ul\u0026gt;\n\u0026lt;h3\u0026gt;BASIC QUALIFICATIONS:\u0026lt;/h3\u0026gt;\n\u0026lt;ul\u0026gt;\n\u0026lt;li\u0026gt;Deep low-level systems programming (C/C++ or Rust)\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Experience with large-scale, high-concurrent production serving.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Strong background in system optimizations: batching, caching, load balancing, parallelism.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Low-level inference optimizations: GPU kernels, code generation.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Experience with testing, benchmarking, and reliability of inference services.\u0026lt;/li\u0026gt;\n\u0026lt;li\u0026gt;Experience designing and implementing CI/CD infrastructure for inference.\u0026lt;/li\u0026gt;\n\u0026lt;/ul\u0026gt;\n\u0026lt;h3\u0026gt;COMPENSATION AND BENEFITS:\u0026lt;/h3\u0026gt;\n\u0026lt;p\u0026gt;$180,000 - $440,000 USD\u0026lt;/p\u0026gt;\n\u0026lt;p\u0026gt;Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short \u0026amp;amp; long-term disability insurance, life insurance, and various other discounts and perks.\u0026lt;/p\u0026gt;\u0026lt;div class=\u0026quot;content-conclusion\u0026quot;\u0026gt;\u0026lt;p\u0026gt;\u0026lt;em\u0026gt;SpaceXAI is an equal opportunity employer. For details on data processing, view our \u0026lt;/em\u0026gt;\u0026lt;em\u0026gt;\u0026lt;a href=\u0026quot;https://x.ai/legal/recruitment-privacy-notice\u0026quot; target=\u0026quot;_blank\u0026quot;\u0026gt;Recruitment Privacy Notice\u0026lt;/a\u0026gt;.\u0026lt;/em\u0026gt;\u0026lt;/p\u0026gt;\u0026lt;/div\u0026gt;","departments":[{"id":4052172007,"name":"Infrastructure","child_ids":[],"parent_id":4024733007}],"offices":[{"id":4035106007,"name":"Palo Alto, CA","location":"Palo Alto, California","child_ids":[],"parent_id":4054926007}],"ai_disclaimer":null,"include_ai_disclaimer":null,"ai_opt_out_request_url":null}