{"jobs":[{"id":"fc28c0b4-3032-4817-ab9b-389057a924c8","title":"Filmmaker / Storyteller","department":"GTM","team":"GTM","employmentType":"FullTime","location":"San Francisco","shouldDisplayCompensationOnJobPostings":false,"secondaryLocations":[],"publishedAt":"2025-05-29T17:59:25.230+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inference/fc28c0b4-3032-4817-ab9b-389057a924c8","applyUrl":"https://jobs.ashbyhq.com/inference/fc28c0b4-3032-4817-ab9b-389057a924c8/application","descriptionHtml":"<h1><strong>Filmmaker / Storyteller</strong></h1><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> is seeking a Filmmaker / Storyteller to join our team and help define the narrative of building the world's largest distributed GPU cluster. This role combines creative vision with eye-catching content production, crafting stories that capture the magic of what we're building while shipping content that resonates with our users. If you live and breathe video content and can find compelling narratives in complex technical work, we want to hear from you!<br /></p><h2><strong>About Inference.net</strong></h2><p style=\"min-height:1.5em\">We are building a real-time marketplace for AI inference that matches spare GPU capacity inside data centers with demand from developers building AI-powered applications. We currently operate the world's largest distributed GPU cluster, with over 5,000 GPUs, hundreds of individual operators, and millions of gigabytes of VRAM connected to the network at any given moment.</p><p style=\"min-height:1.5em\">We are a small, well-funded team working on difficult, high-impact problems at the intersection of AI and distributed systems. We primarily work in-person from our office in downtown San Francisco. Our investors include A16z CSX and Multicoin. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do.</p><p style=\"min-height:1.5em\"></p><h2><strong>About the Role</strong></h2><p style=\"min-height:1.5em\">As our in-house filmmaker, you'll be embedded with our engineering team, capturing the journey of next-generation AI infrastructure. You'll own our entire video content strategy from ideation to final cut, shipping weekly content that attracts eyeballs while authentically representing our mission. This is a high-output role that demands both creative excellence and experimental mindset.</p><p style=\"min-height:1.5em\"></p><h2><strong>Key Responsibilities</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\"><strong>Content Creation &amp; Production:</strong> Produce 4+ short-form videos monthly, 1 commercial monthly, and 1 documentary quarterly, handling everything from concept to final delivery</p></li><li><p style=\"min-height:1.5em\"><strong>Narrative Development:</strong> Extract compelling stories from our team and technology, making complex AI/distributed systems concepts accessible and engaging</p></li><li><p style=\"min-height:1.5em\"><strong>Platform Strategy:</strong> Ship content weekly across X, YouTube, TikTok, LinkedIn, and emerging platforms, optimizing for each platform's unique audience</p></li><li><p style=\"min-height:1.5em\"><strong>Creative Experimentation:</strong> Constantly test new formats, styles, and approaches to find what resonates with developers and tech enthusiasts</p></li><li><p style=\"min-height:1.5em\"><strong>Brand Storytelling:</strong> Define and evolve our company's visual narrative as we scale from startup to industry leader</p></li><li><p style=\"min-height:1.5em\"><strong>Team Building:</strong> After establishing our content foundation, recruit and manage freelancers, editors, and production talent to scale output<br /></p><p style=\"min-height:1.5em\"></p></li></ul><h2><strong>What We're Looking For</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\"><strong>Portfolio of Excellence:</strong> Demonstrated ability to create viral, high-quality video content across multiple formats</p></li><li><p style=\"min-height:1.5em\"><strong>Storytelling Mastery:</strong> Exceptional ability to find and craft narratives, especially from technical or complex subject matter</p></li><li><p style=\"min-height:1.5em\"><strong>Technical Production Skills:</strong> End-to-end video production capabilities including shooting, editing, motion graphics, and color grading</p></li><li><p style=\"min-height:1.5em\"><strong>Platform Native:</strong> Deep understanding of what performs on modern platforms, from TikTok trends to YouTube optimization</p></li><li><p style=\"min-height:1.5em\"><strong>High Velocity Mindset:</strong> Comfort shipping content weekly while maintaining quality standards</p></li><li><p style=\"min-height:1.5em\"><strong>Collaborative Spirit:</strong> Ability to work closely with engineers and extract authentic stories from technical experts</p></li><li><p style=\"min-height:1.5em\"><strong>Startup DNA:</strong> Thrives in ambiguous, fast-moving environments with changing priorities</p></li><li><p style=\"min-height:1.5em\"><strong>On-Site Commitment:</strong> Available to work in-person from our SF office 5 days per week<br /></p><p style=\"min-height:1.5em\"></p></li></ul><h2><strong>Nice to Have</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience creating content for developer or B2B tech audiences</p></li><li><p style=\"min-height:1.5em\">Background in documentary filmmaking or journalism</p></li><li><p style=\"min-height:1.5em\">Motion graphics and animation skills</p></li><li><p style=\"min-height:1.5em\">Experience managing creative teams or freelancers</p></li></ul><p style=\"min-height:1.5em\"></p><h2><strong>What You'll Create</strong></h2><p style=\"min-height:1.5em\">Your work will span from punchy 30-second clips that stop the scroll to thoughtful mini-documentaries about the future of AI infrastructure. Think less corporate video, more cinematic storytelling that happens to feature GPUs and distributed systems. You'll make content that developers share because it's genuinely good, not just informative.</p><p style=\"min-height:1.5em\"></p><h2><strong>Compensation</strong></h2><p style=\"min-height:1.5em\">We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $100,000 - $140,000, plus competitive equity and benefits including:</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Full healthcare coverage</p></li><li><p style=\"min-height:1.5em\">Quarterly offsites</p></li><li><p style=\"min-height:1.5em\">Flexible PTO</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\">If you're ready to own the visual narrative of AI's next chapter and can ship content that makes infrastructure feel like magic, we'd love to see your work. Please send your portfolio, resume, and a brief note about your favorite piece of content you've created to hiring@inference.net.</p>","descriptionPlain":"FILMMAKER / STORYTELLER\n\nInference.net http://Inference.net is seeking a Filmmaker / Storyteller to join our team and help define the narrative of building the world's largest distributed GPU cluster. This role combines creative vision with eye-catching content production, crafting stories that capture the magic of what we're building while shipping content that resonates with our users. If you live and breathe video content and can find compelling narratives in complex technical work, we want to hear from you!\n\n\n\nABOUT INFERENCE.NET\n\nWe are building a real-time marketplace for AI inference that matches spare GPU capacity inside data centers with demand from developers building AI-powered applications. We currently operate the world's largest distributed GPU cluster, with over 5,000 GPUs, hundreds of individual operators, and millions of gigabytes of VRAM connected to the network at any given moment.\n\nWe are a small, well-funded team working on difficult, high-impact problems at the intersection of AI and distributed systems. We primarily work in-person from our office in downtown San Francisco. Our investors include A16z CSX and Multicoin. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do.\n\n\n\n\nABOUT THE ROLE\n\nAs our in-house filmmaker, you'll be embedded with our engineering team, capturing the journey of next-generation AI infrastructure. You'll own our entire video content strategy from ideation to final cut, shipping weekly content that attracts eyeballs while authentically representing our mission. This is a high-output role that demands both creative excellence and experimental mindset.\n\n\n\n\nKEY RESPONSIBILITIES\n\n - Content Creation & Production: Produce 4+ short-form videos monthly, 1 commercial monthly, and 1 documentary quarterly, handling everything from concept to final delivery\n\n - Narrative Development: Extract compelling stories from our team and technology, making complex AI/distributed systems concepts accessible and engaging\n\n - Platform Strategy: Ship content weekly across X, YouTube, TikTok, LinkedIn, and emerging platforms, optimizing for each platform's unique audience\n\n - Creative Experimentation: Constantly test new formats, styles, and approaches to find what resonates with developers and tech enthusiasts\n\n - Brand Storytelling: Define and evolve our company's visual narrative as we scale from startup to industry leader\n\n - Team Building: After establishing our content foundation, recruit and manage freelancers, editors, and production talent to scale output\n   \n   \n   \n\n\nWHAT WE'RE LOOKING FOR\n\n - Portfolio of Excellence: Demonstrated ability to create viral, high-quality video content across multiple formats\n\n - Storytelling Mastery: Exceptional ability to find and craft narratives, especially from technical or complex subject matter\n\n - Technical Production Skills: End-to-end video production capabilities including shooting, editing, motion graphics, and color grading\n\n - Platform Native: Deep understanding of what performs on modern platforms, from TikTok trends to YouTube optimization\n\n - High Velocity Mindset: Comfort shipping content weekly while maintaining quality standards\n\n - Collaborative Spirit: Ability to work closely with engineers and extract authentic stories from technical experts\n\n - Startup DNA: Thrives in ambiguous, fast-moving environments with changing priorities\n\n - On-Site Commitment: Available to work in-person from our SF office 5 days per week\n   \n   \n   \n\n\nNICE TO HAVE\n\n - Experience creating content for developer or B2B tech audiences\n\n - Background in documentary filmmaking or journalism\n\n - Motion graphics and animation skills\n\n - Experience managing creative teams or freelancers\n\n\n\n\nWHAT YOU'LL CREATE\n\nYour work will span from punchy 30-second clips that stop the scroll to thoughtful mini-documentaries about the future of AI infrastructure. Think less corporate video, more cinematic storytelling that happens to feature GPUs and distributed systems. You'll make content that developers share because it's genuinely good, not just informative.\n\n\n\n\nCOMPENSATION\n\nWe offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $100,000 - $140,000, plus competitive equity and benefits including:\n\n - Full healthcare coverage\n\n - Quarterly offsites\n\n - Flexible PTO\n\n\n\nIf you're ready to own the visual narrative of AI's next chapter and can ship content that makes infrastructure feel like magic, we'd love to see your work. Please send your portfolio, resume, and a brief note about your favorite piece of content you've created to hiring@inference.net.","compensation":{"compensationTierSummary":null,"scrapeableCompensationSalarySummary":null,"compensationTiers":[],"summaryComponents":[]}},{"id":"efe67830-9257-499c-a41c-021c9b96b15f","title":"Fullstack Engineer - Frontend Focus","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"San Francisco","shouldDisplayCompensationOnJobPostings":false,"secondaryLocations":[],"publishedAt":"2025-07-23T16:20:12.428+00:00","isListed":true,"isRemote":true,"workplaceType":"Hybrid","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inference/efe67830-9257-499c-a41c-021c9b96b15f","applyUrl":"https://jobs.ashbyhq.com/inference/efe67830-9257-499c-a41c-021c9b96b15f/application","descriptionHtml":"<p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\"><strong>Inference.net</strong></a><strong> is hiring a Senior Full-Stack (Frontend-Focused) Engineer</strong></p><p style=\"min-height:1.5em\">Help us build beautiful, performant web experiences that give users super-powers over our globally distributed LLM inference platform. If you love shipping React apps that feel snappy at planet-scale, we’d love to meet you.</p><p style=\"min-height:1.5em\"></p><h2><strong>About </strong><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\"><strong>Inference.net</strong></a></h2><p style=\"min-height:1.5em\">We combine idle GPU capacity from around the world into a single cohesive plane of compute capable of serving models like DeepSeek and Llama 4. At any moment, 5,000+ GPUs and hundreds of terabytes of VRAM are connected to our network.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\">We’re a small, well-funded team working in-person from downtown San Francisco (hybrid flexibility as needed). Investors include <strong>a16z CSX</strong> and <strong>Multicoin</strong>. We’re high-agency, collaborative, and obsessed with craft—whether that’s a distributed scheduler or a pixel-perfect UI.</p><p style=\"min-height:1.5em\"></p><h2><strong>What you’ll do</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\"><strong>Own the user experience</strong> – design, build, and polish the dashboards, consoles, and customer-facing apps that let users observe, configure, and pay for inference at scale.</p></li><li><p style=\"min-height:1.5em\"><strong>Ship end-to-end features</strong> – from Figma wireframe to React component to backend API and database migration.</p></li><li><p style=\"min-height:1.5em\"><strong>Design a component system</strong> in React + Tailwind that supports rapid iteration and a cohesive design language.</p></li><li><p style=\"min-height:1.5em\"><strong>Optimize performance</strong> – SSR, code-splitting, hydration, and WebSocket-driven real-time updates that hold up under millions of requests per day.</p></li><li><p style=\"min-height:1.5em\"><strong>Collaborate across disciplines</strong> – work shoulder-to-shoulder with distributed-systems engineers, product designers, and founders to turn complex infrastructure into delightful product.</p></li><li><p style=\"min-height:1.5em\"><strong>Level-up the team</strong> – lead design reviews, mentor junior engineers, and introduce best practices for testing, accessibility, and observability.</p></li></ul><p style=\"min-height:1.5em\"></p><h2><strong>What we’re looking for</strong></h2><p style=\"min-height:1.5em\"><strong>Must-Have</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">5+ years building production React applications</p></li><li><p style=\"min-height:1.5em\">Deep knowledge of Tailwind CSS &amp; modern CSS architecture</p></li><li><p style=\"min-height:1.5em\">Typescript mastery and strong fundamentals in JS/DOM/APIs</p></li><li><p style=\"min-height:1.5em\">Experience designing REST/JSON or gRPC backends (Node, Go, or similar)</p></li><li><p style=\"min-height:1.5em\">AuthN/AuthZ design (OIDC, JWT)</p></li><li><p style=\"min-height:1.5em\">Product sense: you care about UX details &amp; accessibility</p></li></ul><p style=\"min-height:1.5em\"><strong>Nice-to-Have</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience with Tanstack / Next.js</p></li><li><p style=\"min-height:1.5em\">Data-viz libraries (Recharts, Visx, D3)</p></li><li><p style=\"min-height:1.5em\">tRPC experience</p></li><li><p style=\"min-height:1.5em\">Familiarity with GPU or ML tooling dashboards</p></li><li><p style=\"min-height:1.5em\">Comfort debugging perf issues (Lighthouse, Chrome DevTools)</p></li><li><p style=\"min-height:1.5em\">Dev-ops chops: CI/CD, Docker, Terraform</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\">You don’t need to tick every “nice-to-have” box—curiosity and the ability to learn quickly matter more.</p><p style=\"min-height:1.5em\"></p><h2><strong>Compensation</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\"><strong>Base salary:</strong> $120,000 – $180,000</p></li><li><p style=\"min-height:1.5em\"><strong>Equity:</strong> significant early-stage grant</p></li><li><p style=\"min-height:1.5em\"><strong>Benefits:</strong> full medical/dental/vision, 401(k) with match, generous PTO, commuter + hardware stipends, daily office lunch</p></li></ul><p style=\"min-height:1.5em\"></p><h2><strong>How we work</strong></h2><p style=\"min-height:1.5em\">We iterate fast, test in prod (safely!), and celebrate small wins. You’ll demo work twice a week, pair with systems engineers, and ship to users continuously. Most of us are in the office 3–4 days a week; remote candidates considered if time-zone compatible with Pacific hours.</p><p style=\"min-height:1.5em\"></p><h2><strong>Equal Opportunity</strong></h2><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> is an equal opportunity employer. We value diversity and do not discriminate on the basis of race, color, religion, gender identity, sexual orientation, national origin, veteran status, disability, age, or any other protected status.</p><p style=\"min-height:1.5em\"></p><h2><strong>Ready to build the front door to planet-scale AI?</strong></h2><p style=\"min-height:1.5em\">Send a short note and a link to something you’ve shipped (code, demo, or Dribbble shots) to <strong>jobs@inference.net</strong>. We can’t wait to chat!</p>","descriptionPlain":"Inference.net http://Inference.net is hiring a Senior Full-Stack (Frontend-Focused) Engineer\n\nHelp us build beautiful, performant web experiences that give users super-powers over our globally distributed LLM inference platform. If you love shipping React apps that feel snappy at planet-scale, we’d love to meet you.\n\n\n\n\nABOUT INFERENCE.NET http://Inference.net\n\nWe combine idle GPU capacity from around the world into a single cohesive plane of compute capable of serving models like DeepSeek and Llama 4. At any moment, 5,000+ GPUs and hundreds of terabytes of VRAM are connected to our network.\n\n\n\nWe’re a small, well-funded team working in-person from downtown San Francisco (hybrid flexibility as needed). Investors include a16z CSX and Multicoin. We’re high-agency, collaborative, and obsessed with craft—whether that’s a distributed scheduler or a pixel-perfect UI.\n\n\n\n\nWHAT YOU’LL DO\n\n - Own the user experience – design, build, and polish the dashboards, consoles, and customer-facing apps that let users observe, configure, and pay for inference at scale.\n\n - Ship end-to-end features – from Figma wireframe to React component to backend API and database migration.\n\n - Design a component system in React + Tailwind that supports rapid iteration and a cohesive design language.\n\n - Optimize performance – SSR, code-splitting, hydration, and WebSocket-driven real-time updates that hold up under millions of requests per day.\n\n - Collaborate across disciplines – work shoulder-to-shoulder with distributed-systems engineers, product designers, and founders to turn complex infrastructure into delightful product.\n\n - Level-up the team – lead design reviews, mentor junior engineers, and introduce best practices for testing, accessibility, and observability.\n\n\n\n\nWHAT WE’RE LOOKING FOR\n\nMust-Have\n\n - 5+ years building production React applications\n\n - Deep knowledge of Tailwind CSS & modern CSS architecture\n\n - Typescript mastery and strong fundamentals in JS/DOM/APIs\n\n - Experience designing REST/JSON or gRPC backends (Node, Go, or similar)\n\n - AuthN/AuthZ design (OIDC, JWT)\n\n - Product sense: you care about UX details & accessibility\n\nNice-to-Have\n\n - Experience with Tanstack / Next.js\n\n - Data-viz libraries (Recharts, Visx, D3)\n\n - tRPC experience\n\n - Familiarity with GPU or ML tooling dashboards\n\n - Comfort debugging perf issues (Lighthouse, Chrome DevTools)\n\n - Dev-ops chops: CI/CD, Docker, Terraform\n\n\n\nYou don’t need to tick every “nice-to-have” box—curiosity and the ability to learn quickly matter more.\n\n\n\n\nCOMPENSATION\n\n - Base salary: $120,000 – $180,000\n\n - Equity: significant early-stage grant\n\n - Benefits: full medical/dental/vision, 401(k) with match, generous PTO, commuter + hardware stipends, daily office lunch\n\n\n\n\nHOW WE WORK\n\nWe iterate fast, test in prod (safely!), and celebrate small wins. You’ll demo work twice a week, pair with systems engineers, and ship to users continuously. Most of us are in the office 3–4 days a week; remote candidates considered if time-zone compatible with Pacific hours.\n\n\n\n\nEQUAL OPPORTUNITY\n\nInference.net http://Inference.net is an equal opportunity employer. We value diversity and do not discriminate on the basis of race, color, religion, gender identity, sexual orientation, national origin, veteran status, disability, age, or any other protected status.\n\n\n\n\nREADY TO BUILD THE FRONT DOOR TO PLANET-SCALE AI?\n\nSend a short note and a link to something you’ve shipped (code, demo, or Dribbble shots) to jobs@inference.net. We can’t wait to chat!","compensation":{"compensationTierSummary":null,"scrapeableCompensationSalarySummary":null,"compensationTiers":[],"summaryComponents":[]}},{"id":"fba2457b-6ef7-4b65-8086-8bf61790279d","title":"Machine Learning Researcher","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"San Francisco","shouldDisplayCompensationOnJobPostings":true,"secondaryLocations":[],"publishedAt":"2026-01-05T23:52:51.371+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inference/fba2457b-6ef7-4b65-8086-8bf61790279d","applyUrl":"https://jobs.ashbyhq.com/inference/fba2457b-6ef7-4b65-8086-8bf61790279d/application","descriptionHtml":"<p style=\"min-height:1.5em\">Help us push the boundaries of what's possible in LLM post-training. If you love training models, exploring new architectures, running experiments, and turning research insights into products that ship, we'd love to meet you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About </strong><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\"><strong>Inference.net</strong></a></p><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.</p><p style=\"min-height:1.5em\">We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">You will be responsible for conducting research into experimental models, training systems, and modalities to create novel products for our customers. Your work will span from exploring new architectures and learning methods to optimizing latency and efficiency, with the goal of delivering better models to customers.</p><p style=\"min-height:1.5em\">Your north star is pushing the frontier of what's possible in LLM post-training. You'll explore new techniques, run rigorous experiments, and when something works, help bring it into production with the help of your teammates. This includes training models for customers and running evaluations as part of validating your research. This role reports directly to the founding team. You'll have the autonomy, a large compute budget / GPU reservation, and technical support to explore ambitious ideas and ship the ones that work.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>Key Responsibilities</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Research and experiment with new model architectures to improve quality, efficiency, or capability</p></li><li><p style=\"min-height:1.5em\">Explore methods to decrease inference latency and improve serving efficiency</p></li><li><p style=\"min-height:1.5em\">Run experiments with new learning methods, including novel approaches to SFT, RLHF, DPO, and other post-training techniques</p></li><li><p style=\"min-height:1.5em\">Perform reinforcement learning research to improve model alignment and capability</p></li><li><p style=\"min-height:1.5em\">Develop and improve our distillation pipeline for training high-quality models from frontier teachers</p></li><li><p style=\"min-height:1.5em\">Train models for clients and run evaluations to validate research findings in production settings</p></li><li><p style=\"min-height:1.5em\">Create robust benchmarks and evaluation frameworks that ensure custom models match or exceed frontier performance</p></li><li><p style=\"min-height:1.5em\">Stay current with ML research and identify techniques that can improve our platform</p></li><li><p style=\"min-height:1.5em\">Collaborate with applied engineers to bring successful research into production systems</p></li><li><p style=\"min-height:1.5em\">Document findings and share knowledge with the team</p></li></ul><p style=\"min-height:1.5em\"><strong>Requirements</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">3+ years of experience training AI models using PyTorch</p></li><li><p style=\"min-height:1.5em\">Deep understanding of transformer architectures, attention mechanisms, and model internals</p></li><li><p style=\"min-height:1.5em\">Hands-on experience with post-training LLMs using SFT, RLHF, DPO, or other alignment techniques</p></li><li><p style=\"min-height:1.5em\">Experience with LLM-specific training frameworks (e.g., Hugging Face Transformers, DeepSpeed, Megatron, TRL, or similar)</p></li><li><p style=\"min-height:1.5em\">Strong experimental methodology, including ability to design, run, and analyze rigorous experiments</p></li><li><p style=\"min-height:1.5em\">Track record of implementing ideas from recent ML papers</p></li><li><p style=\"min-height:1.5em\">Experience training on NVIDIA GPUs at scale</p></li><li><p style=\"min-height:1.5em\">Strong foundation in ML fundamentals: optimization, loss functions, regularization, generalization</p></li></ul><p style=\"min-height:1.5em\"><strong>Nice-to-Have</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Publications in ML venues</p></li><li><p style=\"min-height:1.5em\">Experience with model distillation or knowledge transfer</p></li><li><p style=\"min-height:1.5em\">Experience with LLM speed optimization techniques</p></li><li><p style=\"min-height:1.5em\">Familiarity with vision encoders, multimodal models, or other modalities</p></li><li><p style=\"min-height:1.5em\">Experience with distributed training and infrastructure at scale</p></li><li><p style=\"min-height:1.5em\">Contributions to open-source ML projects</p></li></ul><p style=\"min-height:1.5em\"><em>You don't need to tick every box. Curiosity and the ability to learn quickly matter more.</em></p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>Compensation</strong></p><p style=\"min-height:1.5em\">We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $250,000 - $350,000, plus equity and benefits, depending on experience.</p><p style=\"min-height:1.5em\"><strong>Equal Opportunity</strong></p><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.</p><p style=\"min-height:1.5em\">If you're excited about pushing the boundaries of custom AI research, we'd love to hear from you. Please send your resume and GitHub to <strong>amar@inference.net</strong> and/or here on Ashby.</p>","descriptionPlain":"Help us push the boundaries of what's possible in LLM post-training. If you love training models, exploring new architectures, running experiments, and turning research insights into products that ship, we'd love to meet you.\n\n\n\nAbout Inference.net http://Inference.net\n\nInference.net http://Inference.net trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.\n\nWe are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.\n\n\n\nAbout the Role\n\nYou will be responsible for conducting research into experimental models, training systems, and modalities to create novel products for our customers. Your work will span from exploring new architectures and learning methods to optimizing latency and efficiency, with the goal of delivering better models to customers.\n\nYour north star is pushing the frontier of what's possible in LLM post-training. You'll explore new techniques, run rigorous experiments, and when something works, help bring it into production with the help of your teammates. This includes training models for customers and running evaluations as part of validating your research. This role reports directly to the founding team. You'll have the autonomy, a large compute budget / GPU reservation, and technical support to explore ambitious ideas and ship the ones that work.\n\n\n\nKey Responsibilities\n\n - Research and experiment with new model architectures to improve quality, efficiency, or capability\n\n - Explore methods to decrease inference latency and improve serving efficiency\n\n - Run experiments with new learning methods, including novel approaches to SFT, RLHF, DPO, and other post-training techniques\n\n - Perform reinforcement learning research to improve model alignment and capability\n\n - Develop and improve our distillation pipeline for training high-quality models from frontier teachers\n\n - Train models for clients and run evaluations to validate research findings in production settings\n\n - Create robust benchmarks and evaluation frameworks that ensure custom models match or exceed frontier performance\n\n - Stay current with ML research and identify techniques that can improve our platform\n\n - Collaborate with applied engineers to bring successful research into production systems\n\n - Document findings and share knowledge with the team\n\nRequirements\n\n - 3+ years of experience training AI models using PyTorch\n\n - Deep understanding of transformer architectures, attention mechanisms, and model internals\n\n - Hands-on experience with post-training LLMs using SFT, RLHF, DPO, or other alignment techniques\n\n - Experience with LLM-specific training frameworks (e.g., Hugging Face Transformers, DeepSpeed, Megatron, TRL, or similar)\n\n - Strong experimental methodology, including ability to design, run, and analyze rigorous experiments\n\n - Track record of implementing ideas from recent ML papers\n\n - Experience training on NVIDIA GPUs at scale\n\n - Strong foundation in ML fundamentals: optimization, loss functions, regularization, generalization\n\nNice-to-Have\n\n - Publications in ML venues\n\n - Experience with model distillation or knowledge transfer\n\n - Experience with LLM speed optimization techniques\n\n - Familiarity with vision encoders, multimodal models, or other modalities\n\n - Experience with distributed training and infrastructure at scale\n\n - Contributions to open-source ML projects\n\nYou don't need to tick every box. Curiosity and the ability to learn quickly matter more.\n\n\n\nCompensation\n\nWe offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $250,000 - $350,000, plus equity and benefits, depending on experience.\n\nEqual Opportunity\n\nInference.net http://Inference.net is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.\n\nIf you're excited about pushing the boundaries of custom AI research, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or here on Ashby.","compensation":{"compensationTierSummary":"$250K – $350K • Offers Equity","scrapeableCompensationSalarySummary":"$250K - $350K","compensationTiers":[{"id":"791f1669-b14a-480a-bff0-262dbbd0681f","tierSummary":"$250K – $350K • Offers Equity","title":null,"additionalInformation":null,"components":[{"id":"86d22a16-03f6-4496-be3b-bd39c73737f4","summary":"Offers Equity","compensationType":"EquityPercentage","interval":"NONE","currencyCode":null,"minValue":null,"maxValue":null},{"id":"ca0b7bb1-a97e-4e12-a017-647bd468f852","summary":"$250K – $350K","compensationType":"Salary","interval":"1 YEAR","currencyCode":"USD","minValue":250000,"maxValue":350000}]}],"summaryComponents":[{"compensationType":"EquityPercentage","interval":"NONE","currencyCode":null,"minValue":null,"maxValue":null},{"compensationType":"Salary","interval":"1 YEAR","currencyCode":"USD","minValue":250000,"maxValue":350000}]}},{"id":"b6798110-1218-4ac0-bcf4-9ee8b51810cb","title":"Applied Machine Learning Engineer","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"San Francisco","shouldDisplayCompensationOnJobPostings":true,"secondaryLocations":[],"publishedAt":"2026-01-05T23:52:39.424+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inference/b6798110-1218-4ac0-bcf4-9ee8b51810cb","applyUrl":"https://jobs.ashbyhq.com/inference/b6798110-1218-4ac0-bcf4-9ee8b51810cb/application","descriptionHtml":"<p style=\"min-height:1.5em\">Help us build the systems that train specialized AI models for the fastest-growing companies in the world. If you love taking cutting-edge ML techniques and turning them into products that ship, we'd love to meet you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About </strong><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\"><strong>Inference.net</strong></a></p><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.</p><p style=\"min-height:1.5em\">We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">You will be responsible for building and improving the core ML systems that power our custom model training platform, while also applying these systems directly for customers. Your role sits at the intersection of applied research and production engineering. You'll lead projects from data intake to trained model, building the infrastructure and tooling along the way.</p><p style=\"min-height:1.5em\">Your north star is model quality at scale, measured by how well our custom models match frontier performance, how efficiently we can train and serve them, and how smoothly we can deliver results to our customers. You'll own the full training lifecycle: processing data, creating dashboards for visibility, training models using our frameworks, running evaluations, and shipping results. This role reports directly to the founding team. You'll have the autonomy, a large compute budget / GPU reservation, and technical support to push the boundaries of what's possible in custom model training.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>Key Responsibilities</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Lead projects from from data intake through the full training pipeline, including processing, cleaning, and preparing datasets for model training</p></li><li><p style=\"min-height:1.5em\">Build and maintain data processing pipelines for aggregating, transforming, and validating training data</p></li><li><p style=\"min-height:1.5em\">Create dashboards and visualization tools to display training metrics, data quality, and model performance</p></li><li><p style=\"min-height:1.5em\">Train models using our internal frameworks and iterate based on evaluation results</p></li><li><p style=\"min-height:1.5em\">Develop robust benchmarks and evaluation frameworks that ensure custom models match or exceed frontier performance</p></li><li><p style=\"min-height:1.5em\">Build systems to automate portions of the training workflow, reducing manual intervention and improving consistency</p></li><li><p style=\"min-height:1.5em\">Take research features and ship them into production settings</p></li><li><p style=\"min-height:1.5em\">Apply the latest techniques in SFT, RL, and model optimization to improve training quality and efficiency</p></li><li><p style=\"min-height:1.5em\">Collaborate with infrastructure engineers to scale training across our GPU fleet</p></li><li><p style=\"min-height:1.5em\">Deeply understand customer use cases to inform training strategies and surface edge cases</p></li></ul><p style=\"min-height:1.5em\"><strong>Requirements</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">2+ years of experience training AI models using PyTorch</p></li><li><p style=\"min-height:1.5em\">Hands-on experience with post-training LLMs using SFT or RL</p></li><li><p style=\"min-height:1.5em\">Strong understanding of transformer architectures and how they're trained</p></li><li><p style=\"min-height:1.5em\">Experience with LLM-specific training frameworks (e.g., Hugging Face Transformers, DeepSpeed, Axolotl, or similar)</p></li><li><p style=\"min-height:1.5em\">Experience training on NVIDIA GPUs</p></li><li><p style=\"min-height:1.5em\">Strong data processing skills and comfortable building ETL pipelines and working with large datasets</p></li><li><p style=\"min-height:1.5em\">Track record of creating benchmarks and evaluations</p></li><li><p style=\"min-height:1.5em\">Ability to take research techniques and apply them to production systems </p></li></ul><p style=\"min-height:1.5em\"><strong>Nice-to-Have</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience with model distillation or knowledge transfer</p></li><li><p style=\"min-height:1.5em\">Experience building dashboards and data visualization tools</p></li><li><p style=\"min-height:1.5em\">Familiarity with vision encoders and multimodal models</p></li><li><p style=\"min-height:1.5em\">Experience with distributed training at scale</p></li><li><p style=\"min-height:1.5em\">Contributions to open-source ML projects</p></li></ul><p style=\"min-height:1.5em\"><em>You don't need to tick every box. Curiosity and the ability to learn quickly matter more.</em></p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>Compensation</strong></p><p style=\"min-height:1.5em\">We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.</p><p style=\"min-height:1.5em\"><strong>Equal Opportunity</strong></p><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.</p><p style=\"min-height:1.5em\">If you're excited about building the future of custom AI infrastructure, we'd love to hear from you. Please send your resume and GitHub to <strong>amar@inference.net</strong> and/or apply here on Ashby.</p>","descriptionPlain":"Help us build the systems that train specialized AI models for the fastest-growing companies in the world. If you love taking cutting-edge ML techniques and turning them into products that ship, we'd love to meet you.\n\n\n\nAbout Inference.net http://Inference.net\n\nInference.net http://Inference.net trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.\n\nWe are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.\n\n\n\nAbout the Role\n\nYou will be responsible for building and improving the core ML systems that power our custom model training platform, while also applying these systems directly for customers. Your role sits at the intersection of applied research and production engineering. You'll lead projects from data intake to trained model, building the infrastructure and tooling along the way.\n\nYour north star is model quality at scale, measured by how well our custom models match frontier performance, how efficiently we can train and serve them, and how smoothly we can deliver results to our customers. You'll own the full training lifecycle: processing data, creating dashboards for visibility, training models using our frameworks, running evaluations, and shipping results. This role reports directly to the founding team. You'll have the autonomy, a large compute budget / GPU reservation, and technical support to push the boundaries of what's possible in custom model training.\n\n\n\nKey Responsibilities\n\n - Lead projects from from data intake through the full training pipeline, including processing, cleaning, and preparing datasets for model training\n\n - Build and maintain data processing pipelines for aggregating, transforming, and validating training data\n\n - Create dashboards and visualization tools to display training metrics, data quality, and model performance\n\n - Train models using our internal frameworks and iterate based on evaluation results\n\n - Develop robust benchmarks and evaluation frameworks that ensure custom models match or exceed frontier performance\n\n - Build systems to automate portions of the training workflow, reducing manual intervention and improving consistency\n\n - Take research features and ship them into production settings\n\n - Apply the latest techniques in SFT, RL, and model optimization to improve training quality and efficiency\n\n - Collaborate with infrastructure engineers to scale training across our GPU fleet\n\n - Deeply understand customer use cases to inform training strategies and surface edge cases\n\nRequirements\n\n - 2+ years of experience training AI models using PyTorch\n\n - Hands-on experience with post-training LLMs using SFT or RL\n\n - Strong understanding of transformer architectures and how they're trained\n\n - Experience with LLM-specific training frameworks (e.g., Hugging Face Transformers, DeepSpeed, Axolotl, or similar)\n\n - Experience training on NVIDIA GPUs\n\n - Strong data processing skills and comfortable building ETL pipelines and working with large datasets\n\n - Track record of creating benchmarks and evaluations\n\n - Ability to take research techniques and apply them to production systems \n\nNice-to-Have\n\n - Experience with model distillation or knowledge transfer\n\n - Experience building dashboards and data visualization tools\n\n - Familiarity with vision encoders and multimodal models\n\n - Experience with distributed training at scale\n\n - Contributions to open-source ML projects\n\nYou don't need to tick every box. Curiosity and the ability to learn quickly matter more.\n\n\n\nCompensation\n\nWe offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.\n\nEqual Opportunity\n\nInference.net http://Inference.net is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.\n\nIf you're excited about building the future of custom AI infrastructure, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or apply here on Ashby.","compensation":{"compensationTierSummary":"$220K – $320K • Offers Equity","scrapeableCompensationSalarySummary":"$220K - $320K","compensationTiers":[{"id":"40e80cdf-16ee-4c4b-b7a1-4198b69c1f50","tierSummary":"Estimated Base Salary $220K – $320K • Offers Equity","title":null,"additionalInformation":null,"components":[{"id":"a86e0543-9ee8-4a9d-b2d2-5b102e160c0f","summary":"Estimated Base Salary $220K – $320K","compensationType":"Salary","interval":"1 YEAR","currencyCode":"USD","minValue":220000,"maxValue":320000},{"id":"3147d3e4-8e56-46d7-92c2-56d4e80225af","summary":"Offers Equity","compensationType":"EquityPercentage","interval":"NONE","currencyCode":null,"minValue":null,"maxValue":null}]}],"summaryComponents":[{"compensationType":"Salary","interval":"1 YEAR","currencyCode":"USD","minValue":220000,"maxValue":320000},{"compensationType":"EquityPercentage","interval":"NONE","currencyCode":null,"minValue":null,"maxValue":null}]}},{"id":"7a2963de-1b33-4dfc-b711-990faa93a6a5","title":"Senior Software Engineer - Model Performance","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"San Francisco","shouldDisplayCompensationOnJobPostings":false,"secondaryLocations":[],"publishedAt":"2026-01-21T18:33:05.651+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inference/7a2963de-1b33-4dfc-b711-990faa93a6a5","applyUrl":"https://jobs.ashbyhq.com/inference/7a2963de-1b33-4dfc-b711-990faa93a6a5/application","descriptionHtml":"<p style=\"min-height:1.5em\">Help us make inference blazingly fast. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet you.</p><p style=\"min-height:1.5em\"><strong>About </strong><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\"><strong>Inference.net</strong></a></p><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.</p><p style=\"min-height:1.5em\">We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.</p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">You will be responsible for making our inference stack as fast and efficient as possible. Your work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale.</p><p style=\"min-height:1.5em\">Your north star is inference performance: latency, throughput, cost efficiency, and how quickly we can bring new model architectures into production. You'll work across the full inference stack—from CUDA kernels to serving frameworks—to find and eliminate bottlenecks. This role reports directly to the founding team. You'll have the autonomy, a large compute budget, and technical support to push the limits of what's possible in model serving.</p><p style=\"min-height:1.5em\"><strong>Key Responsibilities</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving</p></li><li><p style=\"min-height:1.5em\">Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance</p></li><li><p style=\"min-height:1.5em\">Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure</p></li><li><p style=\"min-height:1.5em\">Add support for new model architectures, ensuring they meet our performance standards before going to production</p></li><li><p style=\"min-height:1.5em\">Experiment with novel inference techniques and bring successful approaches into production</p></li><li><p style=\"min-height:1.5em\">Build tooling and benchmarks to measure and track inference performance across our fleet</p></li><li><p style=\"min-height:1.5em\">Collaborate with applied ML engineers to ensure trained models can be served efficiently</p></li></ul><p style=\"min-height:1.5em\"><strong>Requirements</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">2+ years of experience in ML systems, inference optimization, or GPU programming</p></li><li><p style=\"min-height:1.5em\">Strong proficiency in Python and familiarity with C++</p></li><li><p style=\"min-height:1.5em\">Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)</p></li><li><p style=\"min-height:1.5em\">Deep understanding of GPU architecture and experience profiling GPU workloads</p></li><li><p style=\"min-height:1.5em\">Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)</p></li><li><p style=\"min-height:1.5em\">Experience with PyTorch and understanding of how models execute on hardware</p></li><li><p style=\"min-height:1.5em\">Track record of measurably improving system performance</p></li></ul><p style=\"min-height:1.5em\"><strong>Nice-to-Have</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience with CUDA programming</p></li><li><p style=\"min-height:1.5em\">Familiarity with serving non-LLM models (TTS, vision, embeddings)</p></li><li><p style=\"min-height:1.5em\">Experience with distributed inference and multi-GPU serving</p></li><li><p style=\"min-height:1.5em\">Contributions to open-source inference frameworks</p></li><li><p style=\"min-height:1.5em\">Experience with Docker and Kubernetes</p></li></ul><p style=\"min-height:1.5em\">You don't need to tick every box. Curiosity and the ability to learn quickly matter more.</p><p style=\"min-height:1.5em\"><strong>Compensation</strong></p><p style=\"min-height:1.5em\">We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.</p><p style=\"min-height:1.5em\"><strong>Equal Opportunity</strong></p><p style=\"min-height:1.5em\"><a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://Inference.net\">Inference.net</a> is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.</p><p style=\"min-height:1.5em\">If you're excited about making AI inference faster for everyone, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or apply here on Ashby.</p><p style=\"min-height:1.5em\"></p>","descriptionPlain":"Help us make inference blazingly fast. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet you.\n\nAbout Inference.net http://Inference.net\n\nInference.net http://Inference.net trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.\n\nWe are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.\n\nAbout the Role\n\nYou will be responsible for making our inference stack as fast and efficient as possible. Your work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale.\n\nYour north star is inference performance: latency, throughput, cost efficiency, and how quickly we can bring new model architectures into production. You'll work across the full inference stack—from CUDA kernels to serving frameworks—to find and eliminate bottlenecks. This role reports directly to the founding team. You'll have the autonomy, a large compute budget, and technical support to push the limits of what's possible in model serving.\n\nKey Responsibilities\n\n - Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving\n\n - Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance\n\n - Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure\n\n - Add support for new model architectures, ensuring they meet our performance standards before going to production\n\n - Experiment with novel inference techniques and bring successful approaches into production\n\n - Build tooling and benchmarks to measure and track inference performance across our fleet\n\n - Collaborate with applied ML engineers to ensure trained models can be served efficiently\n\nRequirements\n\n - 2+ years of experience in ML systems, inference optimization, or GPU programming\n\n - Strong proficiency in Python and familiarity with C++\n\n - Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)\n\n - Deep understanding of GPU architecture and experience profiling GPU workloads\n\n - Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)\n\n - Experience with PyTorch and understanding of how models execute on hardware\n\n - Track record of measurably improving system performance\n\nNice-to-Have\n\n - Experience with CUDA programming\n\n - Familiarity with serving non-LLM models (TTS, vision, embeddings)\n\n - Experience with distributed inference and multi-GPU serving\n\n - Contributions to open-source inference frameworks\n\n - Experience with Docker and Kubernetes\n\nYou don't need to tick every box. Curiosity and the ability to learn quickly matter more.\n\nCompensation\n\nWe offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.\n\nEqual Opportunity\n\nInference.net http://Inference.net is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.\n\nIf you're excited about making AI inference faster for everyone, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or apply here on Ashby.\n\n","compensation":{"compensationTierSummary":null,"scrapeableCompensationSalarySummary":null,"compensationTiers":[],"summaryComponents":[]}}],"apiVersion":"1"}