{"jobs":[{"id":"49df3034-21c5-47da-b63f-22bf1ab7b893","title":"Expression of Interest - Intelligent Systems Engineer","department":"Expression of Interest","team":"Expression of Interest","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-02-25T12:57:50.992+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/49df3034-21c5-47da-b63f-22bf1ab7b893","applyUrl":"https://jobs.ashbyhq.com/callosum/49df3034-21c5-47da-b63f-22bf1ab7b893/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><p style=\"min-height:1.5em\"><strong>Open Application</strong></p><p style=\"min-height:1.5em\">We're always looking for exceptional engineers who want to work on the hardest problems in AI infrastructure - bringing novel accelerators to life, building the systems that orchestrate them, and charting the new territories of what heterogeneous compute can do.</p><p style=\"min-height:1.5em\">If you think you have unique skills to contribute, or have  experience in any of the following, we'd love to hear from you:</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">High-performance systems and low-level performance engineering</p></li><li><p style=\"min-height:1.5em\">Inference infrastructure, orchestration, or distributed systems</p></li><li><p style=\"min-height:1.5em\">Hardware bring-up, kernel development, or accelerator programming</p></li><li><p style=\"min-height:1.5em\">Simulation, modelling, or workload design</p></li></ul><p style=\"min-height:1.5em\">This is a general interest form rather than an application for a specific role. We prioritise active applicants for open positions, but we review all submissions on an ongoing basis and will reach out when we see a strong match. If you spot a specific role that excites you, we'd encourage you to apply directly.</p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n \n\nOpen Application\n\nWe're always looking for exceptional engineers who want to work on the hardest problems in AI infrastructure - bringing novel accelerators to life, building the systems that orchestrate them, and charting the new territories of what heterogeneous compute can do.\n\nIf you think you have unique skills to contribute, or have  experience in any of the following, we'd love to hear from you:\n\n - High-performance systems and low-level performance engineering\n\n - Inference infrastructure, orchestration, or distributed systems\n\n - Hardware bring-up, kernel development, or accelerator programming\n\n - Simulation, modelling, or workload design\n\nThis is a general interest form rather than an application for a specific role. We prioritise active applicants for open positions, but we review all submissions on an ongoing basis and will reach out when we see a strong match. If you spot a specific role that excites you, we'd encourage you to apply directly."},{"id":"1dcfaab8-ec0e-4374-9880-a1b65465b458","title":"Inference Engine Development - Member of Technical Staff","department":"Compute & Infrastructure","team":"Compute & Infrastructure","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-05-20T16:36:20.026+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/1dcfaab8-ec0e-4374-9880-a1b65465b458","applyUrl":"https://jobs.ashbyhq.com/callosum/1dcfaab8-ec0e-4374-9880-a1b65465b458/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. We are building that vision: infrastructure that treats the full landscape of compute as a unified, co-evolving system, evolved beyond GPUs.</p><p style=\"min-height:1.5em\">Inference engines were designed for single-model inference on homogeneous GPU clusters - this role builds them beyond that. Working directly on systems like vLLM and SGLang, you will adapt and extend them for heterogeneous resources, making them hardware-aware, with deeper optimisation around scheduling, memory, and execution. The execution strategies you design - parallelism, disaggregation, caching - will define what heterogeneous inference looks like at production scale. Your work ensures that the capabilities exposed by the lower layers of the stack translate into real, measurable gains, the new standard for how inference runs on diverse hardware.</p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Contribute upstream to SGLang and vLLM, and maintain internal forks where our requirements diverge</p></li><li><p style=\"min-height:1.5em\">Improve hardware-awareness within inference engines so that scheduling, memory management, and execution adapt to the capabilities of the underlying accelerator</p></li><li><p style=\"min-height:1.5em\">Design and implement bespoke parallelism and disaggregation strategies that go beyond default configurations to better exploit heterogeneous hardware</p></li><li><p style=\"min-height:1.5em\">Work closely with an Accelerator Systems Software engineer to ensure engine-level abstractions map cleanly onto diverse hardware capabilities</p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Deep familiarity with the internals of SGLang, vLLM, or comparable inference serving frameworks - scheduler design, memory management, and execution pipelines</p></li><li><p style=\"min-height:1.5em\">Strong background in high-performance Python and C++/CUDA systems, particularly in the context of ML inference</p></li><li><p style=\"min-height:1.5em\">Experience designing or implementing parallelism strategies for large model serving</p></li><li><p style=\"min-height:1.5em\">Understanding of disaggregated serving architectures and the tradeoffs involved in separating modules of a workflow</p></li><li><p style=\"min-height:1.5em\">Demonstrable record of working effectively in fast-moving open source codebases with evolving APIs and design conventions</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n \n\nAbout the Role\n\nCallosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. We are building that vision: infrastructure that treats the full landscape of compute as a unified, co-evolving system, evolved beyond GPUs.\n\nInference engines were designed for single-model inference on homogeneous GPU clusters - this role builds them beyond that. Working directly on systems like vLLM and SGLang, you will adapt and extend them for heterogeneous resources, making them hardware-aware, with deeper optimisation around scheduling, memory, and execution. The execution strategies you design - parallelism, disaggregation, caching - will define what heterogeneous inference looks like at production scale. Your work ensures that the capabilities exposed by the lower layers of the stack translate into real, measurable gains, the new standard for how inference runs on diverse hardware.\n\nWhat You'll Build\n\n - Contribute upstream to SGLang and vLLM, and maintain internal forks where our requirements diverge\n\n - Improve hardware-awareness within inference engines so that scheduling, memory management, and execution adapt to the capabilities of the underlying accelerator\n\n - Design and implement bespoke parallelism and disaggregation strategies that go beyond default configurations to better exploit heterogeneous hardware\n\n - Work closely with an Accelerator Systems Software engineer to ensure engine-level abstractions map cleanly onto diverse hardware capabilities\n\nWhat Sets You Apart\n\n - Deep familiarity with the internals of SGLang, vLLM, or comparable inference serving frameworks - scheduler design, memory management, and execution pipelines\n\n - Strong background in high-performance Python and C++/CUDA systems, particularly in the context of ML inference\n\n - Experience designing or implementing parallelism strategies for large model serving\n\n - Understanding of disaggregated serving architectures and the tradeoffs involved in separating modules of a workflow\n\n - Demonstrable record of working effectively in fast-moving open source codebases with evolving APIs and design conventions\n\n\n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"35a8cced-2c23-431a-8198-9dacdaaad9d3","title":"Accelerator Systems Software - Member of Technical Staff","department":"Compute & Infrastructure","team":"Compute & Infrastructure","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-05-20T16:35:53.586+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/35a8cced-2c23-431a-8198-9dacdaaad9d3","applyUrl":"https://jobs.ashbyhq.com/callosum/35a8cced-2c23-431a-8198-9dacdaaad9d3/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. We are building that vision: infrastructure that treats the full landscape of compute as a unified, co-evolving system. Callosum is purposefully placed to be the first place to access and deploy new chips, expanding beyond GPUs to enable a system that works in harmony, greater than the sum of its parts.</p><p style=\"min-height:1.5em\">This role sits at the foundation of Callosum’s stack, enabling the company to run AI workloads beyond the constraints of any single hardware vendor. You will build the low-level systems software - kernels, drivers, runtime tools - that make diverse and novel accelerators viable for real-world inference, surfacing the full strength of new silicon. This infrastructure is what enables us to turn a fragmented accelerator landscape into our platform; you will own the design decisions that directly influence performance and reliability.</p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Build and maintain kernels, device drivers, and firmware for heterogeneous accelerators</p></li><li><p style=\"min-height:1.5em\">Design and implement accelerator scheduling at the execution level — kernel launches, dataflow optimisation, and resource management across diverse hardware</p></li><li><p style=\"min-height:1.5em\">Optimise execution paths for latency, throughput, and resource utilisation across accelerator types</p></li><li><p style=\"min-height:1.5em\">Work closely with internal teams and hardware vendors to onboard new accelerator platforms</p></li><li><p style=\"min-height:1.5em\">Contribute to low-level runtime software that bridges our inference stack and the underlying hardware</p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Demonstrable interest in a variety of accelerator microarchitectures</p></li><li><p style=\"min-height:1.5em\">Deep experience in kernel development, device driver authoring, or firmware engineering</p></li><li><p style=\"min-height:1.5em\">Familiarity with accelerator programming models (CUDA, ROCm, or vendor-specific SDKs)</p></li><li><p style=\"min-height:1.5em\">Experience with compiler infrastructure such as MLIR or LLVM</p></li><li><p style=\"min-height:1.5em\">Strong debugging skills in environments with limited tooling and incomplete documentation</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n \n\nAbout the Role\n\nCallosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. We are building that vision: infrastructure that treats the full landscape of compute as a unified, co-evolving system. Callosum is purposefully placed to be the first place to access and deploy new chips, expanding beyond GPUs to enable a system that works in harmony, greater than the sum of its parts.\n\nThis role sits at the foundation of Callosum’s stack, enabling the company to run AI workloads beyond the constraints of any single hardware vendor. You will build the low-level systems software - kernels, drivers, runtime tools - that make diverse and novel accelerators viable for real-world inference, surfacing the full strength of new silicon. This infrastructure is what enables us to turn a fragmented accelerator landscape into our platform; you will own the design decisions that directly influence performance and reliability.\n\nWhat You'll Build\n\n - Build and maintain kernels, device drivers, and firmware for heterogeneous accelerators\n\n - Design and implement accelerator scheduling at the execution level — kernel launches, dataflow optimisation, and resource management across diverse hardware\n\n - Optimise execution paths for latency, throughput, and resource utilisation across accelerator types\n\n - Work closely with internal teams and hardware vendors to onboard new accelerator platforms\n\n - Contribute to low-level runtime software that bridges our inference stack and the underlying hardware\n\nWhat Sets You Apart\n\n - Demonstrable interest in a variety of accelerator microarchitectures\n\n - Deep experience in kernel development, device driver authoring, or firmware engineering\n\n - Familiarity with accelerator programming models (CUDA, ROCm, or vendor-specific SDKs)\n\n - Experience with compiler infrastructure such as MLIR or LLVM\n\n - Strong debugging skills in environments with limited tooling and incomplete documentation\n\n\n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"2e9b3645-074e-48bf-9edb-c04efdfd95d7","title":"Inference Performance & Deployment - Member of Technical Staff","department":"Compute & Infrastructure","team":"Compute & Infrastructure","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-05-20T16:36:34.130+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/2e9b3645-074e-48bf-9edb-c04efdfd95d7","applyUrl":"https://jobs.ashbyhq.com/callosum/2e9b3645-074e-48bf-9edb-c04efdfd95d7/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. We are building that vision: infrastructure that treats the full landscape of compute as a unified, co-evolving system, evolved beyond GPUs.</p><p style=\"min-height:1.5em\">This role owns the bridge between Callosum's internal engineering and the real world. You design the tooling and methodologies that ground our technology in real-world performance and behaviour, sitting at the integration point of every engineering function. You will be the first to run our heterogeneous infrastructure in production-equivalent conditions, systematically characterising performance, identifying bottlenecks, and driving decisions on production-readiness. Your work ensures that every layer of the stack is guided by empirical evidence rather than assumption.</p><p style=\"min-height:1.5em\"><strong>What You’ll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Run experiments self-hosting models on cloud instances or on-prem across providers and hardware configurations, systematically characterising performance envelopes</p></li><li><p style=\"min-height:1.5em\">Develop and maintain deployment patterns that are reproducible, measurable, and optimised for latency, throughput, and cost</p></li><li><p style=\"min-height:1.5em\">Work at the orchestration and routing software that sits above the inference engine - to improve caching, request scheduling, batching, and resource allocation</p></li><li><p style=\"min-height:1.5em\">Act as the integration point for the other roles: consume new accelerator support, engine features, and infrastructure upgrades – to provide high-quality feedback on bottlenecks, essential capabilities, and guide the stack optimisations</p></li><li><p style=\"min-height:1.5em\">Build and maintain benchmarking harnesses, regression suites, and performance dashboards that give the team a shared view of system health and progress</p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience deploying and benchmarking large model inference in production or production-equivalent environments</p></li><li><p style=\"min-height:1.5em\">Familiarity with multi-node GPU deployments and associated networking/communication stacks</p></li><li><p style=\"min-height:1.5em\">Strong end-to-end performance characterisation skills: able to isolate whether a bottleneck is in the network, the runtime, the memory subsystem, or the model itself</p></li><li><p style=\"min-height:1.5em\">Familiarity with serving frameworks like Dynamo, Triton Inference Server, or similar orchestration layers</p></li><li><p style=\"min-height:1.5em\">Clear communication skills - able to translate performance data into actionable, prioritised feedback for the teams building the underlying systems</p></li><li><p style=\"min-height:1.5em\">A demonstrable disciplined and systematic approach to deployment: reproducibility, measurement methodology, controlled comparisons, etc</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n \n\nAbout the Role\n\nCallosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. We are building that vision: infrastructure that treats the full landscape of compute as a unified, co-evolving system, evolved beyond GPUs.\n\nThis role owns the bridge between Callosum's internal engineering and the real world. You design the tooling and methodologies that ground our technology in real-world performance and behaviour, sitting at the integration point of every engineering function. You will be the first to run our heterogeneous infrastructure in production-equivalent conditions, systematically characterising performance, identifying bottlenecks, and driving decisions on production-readiness. Your work ensures that every layer of the stack is guided by empirical evidence rather than assumption.\n\nWhat You’ll Build\n\n - Run experiments self-hosting models on cloud instances or on-prem across providers and hardware configurations, systematically characterising performance envelopes\n\n - Develop and maintain deployment patterns that are reproducible, measurable, and optimised for latency, throughput, and cost\n\n - Work at the orchestration and routing software that sits above the inference engine - to improve caching, request scheduling, batching, and resource allocation\n\n - Act as the integration point for the other roles: consume new accelerator support, engine features, and infrastructure upgrades – to provide high-quality feedback on bottlenecks, essential capabilities, and guide the stack optimisations\n\n - Build and maintain benchmarking harnesses, regression suites, and performance dashboards that give the team a shared view of system health and progress\n\nWhat Sets You Apart\n\n - Experience deploying and benchmarking large model inference in production or production-equivalent environments\n\n - Familiarity with multi-node GPU deployments and associated networking/communication stacks\n\n - Strong end-to-end performance characterisation skills: able to isolate whether a bottleneck is in the network, the runtime, the memory subsystem, or the model itself\n\n - Familiarity with serving frameworks like Dynamo, Triton Inference Server, or similar orchestration layers\n\n - Clear communication skills - able to translate performance data into actionable, prioritised feedback for the teams building the underlying systems\n\n - A demonstrable disciplined and systematic approach to deployment: reproducibility, measurement methodology, controlled comparisons, etc\n\n\n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"c342204a-6bb8-47ae-a926-049a50e950ae","title":"Agent Runtime & Systems - Member of Technical Staff","department":"General Capabilities","team":"General Capabilities","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:34:20.403+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/c342204a-6bb8-47ae-a926-049a50e950ae","applyUrl":"https://jobs.ashbyhq.com/callosum/c342204a-6bb8-47ae-a926-049a50e950ae/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Most agent frameworks are designed for short-lived workflows, best-effort execution, and a narrow set of tools. Callosum is building for a different world: agents that run for weeks, operate across distributed infrastructure, and interact with stateful systems where failures, retries, and side effects must be handled explicitly. Success in this environment requires an agent runtime built entirely from first principles.</p><p style=\"min-height:1.5em\">Sitting at the heart of our technical mission, this position owns the runtime and programming model for long-running agent systems. Your focus will span durable execution, checkpointing and replay, distributed scheduling, secure tool execution, and end-to-end observability. You will develop the core software that makes agent workflows reliable, debuggable, reproducible, and efficient as we expand our model and tool footprint. This is a high-leverage role tackling complex systems challenges across the entire stack.</p><p style=\"min-height:1.5em\">Runtime, observability tooling, and framework design are three separate specialties. We are glad to hire someone with real depth in a combination of them.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Design and build stateful, asynchronous agent runtimes across distributed environments</p></li><li><p style=\"min-height:1.5em\">Define event histories, checkpoint and resume semantics, deterministic reconstruction, retries, cancellation, timeouts, and recovery from partial failure</p></li><li><p style=\"min-height:1.5em\">Build traces and replay, time-travel debugging, and hierarchical tracing across model calls, tools, handoffs, state changes, and runtime decisions</p></li><li><p style=\"min-height:1.5em\">Develop clear Python APIs, libraries, and DSLs with stable extension points for tools, backends, and new execution models</p></li><li><p style=\"min-height:1.5em\">Design scheduling, placement, backpressure, resource accounting, sandboxing, permissions, and isolation for secure tool execution at scale</p></li><li><p style=\"min-height:1.5em\">Develop tooling for trajectory collection and visualisation, performance profiling, experiment comparison, failure attribution, and regression detection across models, prompts, tools, and runtime versions</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What You'll Bring</strong></p><p style=\"min-height:1.5em\">We are looking for real strength in some of the following rather than coverage of all of it.</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Strong distributed systems background with hands-on experience in durable workflows, actor systems, schedulers, workflow engines, distributed databases, or stream processors</p></li><li><p style=\"min-height:1.5em\">Experience designing a library, framework, runtime, or DSL that other engineers or researchers adopted</p></li><li><p style=\"min-height:1.5em\">Deep Python expertise plus proficiency in a systems language such as Rust, C++, or Go - and the instinct to move between abstraction layers when it matters</p></li><li><p style=\"min-height:1.5em\">Hands-on debugging skills across APIs, runtimes, and distributed infrastructure - with a practical understanding of nondeterminism, side effects, and reproducibility</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience building a production-grade runtime, compiler, workflow engine, distributed training system, database, debugger, or orchestration platform</p></li><li><p style=\"min-height:1.5em\">Meaningful contributions to systems such as Ray, Temporal, Kafka, Erlang/BEAM, PyTorch, JAX, TensorFlow, or comparable projects</p></li><li><p style=\"min-height:1.5em\">Strong API design judgement, open-source collaboration, and disciplined approaches to testing, versioning, compatibility, and reproducibility</p></li><li><p style=\"min-height:1.5em\">Experience in ML systems or large-scale inference and training</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nMost agent frameworks are designed for short-lived workflows, best-effort execution, and a narrow set of tools. Callosum is building for a different world: agents that run for weeks, operate across distributed infrastructure, and interact with stateful systems where failures, retries, and side effects must be handled explicitly. Success in this environment requires an agent runtime built entirely from first principles.\n\nSitting at the heart of our technical mission, this position owns the runtime and programming model for long-running agent systems. Your focus will span durable execution, checkpointing and replay, distributed scheduling, secure tool execution, and end-to-end observability. You will develop the core software that makes agent workflows reliable, debuggable, reproducible, and efficient as we expand our model and tool footprint. This is a high-leverage role tackling complex systems challenges across the entire stack.\n\nRuntime, observability tooling, and framework design are three separate specialties. We are glad to hire someone with real depth in a combination of them.\n\n\n\nWhat You'll Build\n\n - Design and build stateful, asynchronous agent runtimes across distributed environments\n\n - Define event histories, checkpoint and resume semantics, deterministic reconstruction, retries, cancellation, timeouts, and recovery from partial failure\n\n - Build traces and replay, time-travel debugging, and hierarchical tracing across model calls, tools, handoffs, state changes, and runtime decisions\n\n - Develop clear Python APIs, libraries, and DSLs with stable extension points for tools, backends, and new execution models\n\n - Design scheduling, placement, backpressure, resource accounting, sandboxing, permissions, and isolation for secure tool execution at scale\n\n - Develop tooling for trajectory collection and visualisation, performance profiling, experiment comparison, failure attribution, and regression detection across models, prompts, tools, and runtime versions\n   \n   \n\nWhat You'll Bring\n\nWe are looking for real strength in some of the following rather than coverage of all of it.\n\n - Strong distributed systems background with hands-on experience in durable workflows, actor systems, schedulers, workflow engines, distributed databases, or stream processors\n\n - Experience designing a library, framework, runtime, or DSL that other engineers or researchers adopted\n\n - Deep Python expertise plus proficiency in a systems language such as Rust, C++, or Go - and the instinct to move between abstraction layers when it matters\n\n - Hands-on debugging skills across APIs, runtimes, and distributed infrastructure - with a practical understanding of nondeterminism, side effects, and reproducibility\n\n\n\nWhat Sets You Apart\n\n - Experience building a production-grade runtime, compiler, workflow engine, distributed training system, database, debugger, or orchestration platform\n\n - Meaningful contributions to systems such as Ray, Temporal, Kafka, Erlang/BEAM, PyTorch, JAX, TensorFlow, or comparable projects\n\n - Strong API design judgement, open-source collaboration, and disciplined approaches to testing, versioning, compatibility, and reproducibility\n\n - Experience in ML systems or large-scale inference and training\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"4ac29d00-3b3c-4c53-8a1f-1798198f5b30","title":"Expression of Interest - Internships","department":"Expression of Interest","team":"Expression of Interest","employmentType":"Intern","location":"London","secondaryLocations":[],"publishedAt":"2026-09-23T12:26:00.606+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/4ac29d00-3b3c-4c53-8a1f-1798198f5b30","applyUrl":"https://jobs.ashbyhq.com/callosum/4ac29d00-3b3c-4c53-8a1f-1798198f5b30/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><p style=\"min-height:1.5em\"><strong>Internship - Expression of Interest</strong></p><p style=\"min-height:1.5em\">We're always interested in hearing from exceptional students who want to work on the hardest problems in AI infrastructure - bringing novel accelerators to life, building the systems that orchestrate them, and charting the new territories of what heterogeneous compute can do.</p><p style=\"min-height:1.5em\">There's no formal internship programme or fixed cycle here - internships happen when the need and the right person come together. This form is how we keep track of who to reach out to when that happens. If you think you have unique skills to contribute, or are studying or have experience in any of the following, we'd love to hear from you:</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">High-performance systems and low-level performance engineering</p></li><li><p style=\"min-height:1.5em\">Inference infrastructure, orchestration, or distributed systems</p></li><li><p style=\"min-height:1.5em\">Hardware bring-up, kernel development, or accelerator programming</p></li><li><p style=\"min-height:1.5em\">Simulation, modelling, or workload design</p></li></ul><p style=\"min-height:1.5em\">This is a general interest form rather than an application for a specific internship. We review all submissions on an ongoing basis and will reach out when we see a strong match and the right opportunity comes up.</p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n \n\nInternship - Expression of Interest\n\nWe're always interested in hearing from exceptional students who want to work on the hardest problems in AI infrastructure - bringing novel accelerators to life, building the systems that orchestrate them, and charting the new territories of what heterogeneous compute can do.\n\nThere's no formal internship programme or fixed cycle here - internships happen when the need and the right person come together. This form is how we keep track of who to reach out to when that happens. If you think you have unique skills to contribute, or are studying or have experience in any of the following, we'd love to hear from you:\n\n - High-performance systems and low-level performance engineering\n\n - Inference infrastructure, orchestration, or distributed systems\n\n - Hardware bring-up, kernel development, or accelerator programming\n\n - Simulation, modelling, or workload design\n\nThis is a general interest form rather than an application for a specific internship. We review all submissions on an ongoing basis and will reach out when we see a strong match and the right opportunity comes up."},{"id":"8262ffb3-fcc0-4b71-8446-61b4316a4561","title":"Networking & Interconnect Systems - Member of Technical Staff","department":"Compute & Infrastructure","team":"Compute & Infrastructure","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:31:34.086+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/8262ffb3-fcc0-4b71-8446-61b4316a4561","applyUrl":"https://jobs.ashbyhq.com/callosum/8262ffb3-fcc0-4b71-8446-61b4316a4561/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. As workloads become increasingly disaggregated, the performance of the system is determined not only by where computation happens, but by how efficiently state and data can move between the hardware best suited to each part of the workload. Communication latency sets the boundary on how far that disaggregation can go.</p><p style=\"min-height:1.5em\">This role owns that boundary. You will develop the communication systems that allow heterogeneous accelerators to operate as parts of one system, working across protocols, memory movement, networking hardware, and the software layers above them. That means understanding where overhead actually comes from, removing assumptions inherited from more general-purpose workloads, and finding new ways to move data directly and efficiently across vendor boundaries. The remit is intentionally broad: from RDMA paths and programmable NICs to cache transfer formats and emerging interconnect technologies, you will own the technical strategy for making communication less of a constraint on the systems we can build.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Characterise and simplify communication paths between accelerators, tracing latency and unnecessary overhead across networking, runtime, memory, and device layers</p></li><li><p style=\"min-height:1.5em\">Develop faster data movement across heterogeneous hardware, exploring RDMA, programmable NICs, DPUs, switches, and software approaches to cross-device memory management</p></li><li><p style=\"min-height:1.5em\">Build flexible mechanisms for transferring model state across accelerator types and execution strategies, including KV cache connectors across different forms of parallelism</p></li><li><p style=\"min-height:1.5em\">Evaluate existing and emerging scale-up and scale-out fabrics, informing cluster architecture, infrastructure integration, and how heterogeneous accelerator communication should evolve</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Deep understanding of networking and communication systems, with the ability to reason across protocols, hardware, memory systems, and software layers</p></li><li><p style=\"min-height:1.5em\">Experience with high-performance data movement such as RDMA, Ethernet fabrics, accelerator interconnects, programmable networking, or similar latency-sensitive systems</p></li><li><p style=\"min-height:1.5em\">Strong systems performance instincts, with the ability to trace bottlenecks across abstraction layers and distinguish fundamental constraints from accidental ones</p></li><li><p style=\"min-height:1.5em\">A demonstrable tendency to question existing abstractions, generalise across unfamiliar technologies, and simplify systems around the requirements of the workload</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nCallosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. As workloads become increasingly disaggregated, the performance of the system is determined not only by where computation happens, but by how efficiently state and data can move between the hardware best suited to each part of the workload. Communication latency sets the boundary on how far that disaggregation can go.\n\nThis role owns that boundary. You will develop the communication systems that allow heterogeneous accelerators to operate as parts of one system, working across protocols, memory movement, networking hardware, and the software layers above them. That means understanding where overhead actually comes from, removing assumptions inherited from more general-purpose workloads, and finding new ways to move data directly and efficiently across vendor boundaries. The remit is intentionally broad: from RDMA paths and programmable NICs to cache transfer formats and emerging interconnect technologies, you will own the technical strategy for making communication less of a constraint on the systems we can build.\n\n\n\nWhat You'll Build\n\n - Characterise and simplify communication paths between accelerators, tracing latency and unnecessary overhead across networking, runtime, memory, and device layers\n\n - Develop faster data movement across heterogeneous hardware, exploring RDMA, programmable NICs, DPUs, switches, and software approaches to cross-device memory management\n\n - Build flexible mechanisms for transferring model state across accelerator types and execution strategies, including KV cache connectors across different forms of parallelism\n\n - Evaluate existing and emerging scale-up and scale-out fabrics, informing cluster architecture, infrastructure integration, and how heterogeneous accelerator communication should evolve\n   \n   \n\nWhat Sets You Apart\n\n - Deep understanding of networking and communication systems, with the ability to reason across protocols, hardware, memory systems, and software layers\n\n - Experience with high-performance data movement such as RDMA, Ethernet fabrics, accelerator interconnects, programmable networking, or similar latency-sensitive systems\n\n - Strong systems performance instincts, with the ability to trace bottlenecks across abstraction layers and distinguish fundamental constraints from accidental ones\n\n - A demonstrable tendency to question existing abstractions, generalise across unfamiliar technologies, and simplify systems around the requirements of the workload\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"07abc888-1e91-4a04-ba37-f463ae7e864c","title":"Research Engineer, Evals - Member of Technical Staff","department":"General Capabilities","team":"General Capabilities","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:34:51.554+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/07abc888-1e91-4a04-ba37-f463ae7e864c","applyUrl":"https://jobs.ashbyhq.com/callosum/07abc888-1e91-4a04-ba37-f463ae7e864c/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous models and hardware. Our team is developing the science that makes that possible: a principled, evidence-based method for building agentic systems automatically at unprecedented scale.</p><p style=\"min-height:1.5em\">The team tackles these problems on two fronts. We build the tools to analyse and evaluate agentic systems rigorously enough to say where and why they go wrong, and we use what we learn to design better ones. We are not simply building a better harness; the best harness will change with every new model, task and generation of silicon. We are building the layer beneath it, so that design decisions follow from evidence rather than taste.</p><p style=\"min-height:1.5em\">This role owns the first of those fronts. Agentic evaluation today is not up to the job: benchmarks saturate, scores move for reasons unrelated to capability, results fail to reproduce across runs, and a single number tells you nothing about which step went wrong. Every design decision the rest of the team makes rests on the quality of that instrument. You will build measures whose construct we can defend, whose variance we understand, and which localise a failure rather than merely scoring it.</p><p style=\"min-height:1.5em\">The mandate is broad, and most of the questions inside it are still unanswered. You will have wide latitude to choose problems, and your results will shape what the company builds in the future.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Design benchmarks and evaluation suites for agentic behaviour – multi-turn, long-horizon, tool-using, operating in non-stationary environments – including the properties that are hard to score, such as reliability and recovery from error</p></li><li><p style=\"min-height:1.5em\">Build the methodology as well as the harness: statistical power, variance across runs, contamination and construct validity</p></li><li><p style=\"min-height:1.5em\">Be the independent voice on the quality of our own results. Red-team our evaluations, design the sanity checks that catch a flattering number before it reaches a decision.</p></li><li><p style=\"min-height:1.5em\">Turn raw traces into structured evidence. Failure taxonomies, behavioural signatures, and attribution of an outcome to the decision that caused it</p></li><li><p style=\"min-height:1.5em\">Move from measurement to prediction: infer what a system can do from partial evidence, and estimate performance on a task before running it. This rests on defining what each model is genuinely good at, what benchmarks really measure, and what capability profile a real compound task demands</p></li><li><p style=\"min-height:1.5em\">Build evaluation and observability infrastructure as durable instruments the whole company works from</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What You'll Bring</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Evidence that you can run research of your own: you take an open question, design the experiments that settle it, and produce results other people can build on</p></li><li><p style=\"min-height:1.5em\">Deep hands-on experience with LLMs in agentic settings: multi-step tasks, tool use, long horizons, and the specific ways all of it breaks</p></li><li><p style=\"min-height:1.5em\">Statistical discipline. Hypotheses stated in advance, uncertainty quantified, and running comparisons that isolate one variable</p></li><li><p style=\"min-height:1.5em\">Strong engineering. You write the code that runs your experiments and you are comfortable working inside a large shared codebase</p></li><li><p style=\"min-height:1.5em\">Strong communication skills. You are able to turn research findings into clear, prioritised guidance for the teams who will act on them, and to write them up for a wider audience</p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Evaluations or benchmarks you built that other people went on to use</p></li><li><p style=\"min-height:1.5em\">Broad training in empirical method, potentially from a field outside machine learning - experimental design, Bayesian inference, model selection, significance testing, uncertainty quantification</p></li><li><p style=\"min-height:1.5em\">Experience with automated grading and model-based judging, and a rigorous account of its failure modes and its agreement with human raters</p></li><li><p style=\"min-height:1.5em\">A published track record in a relevant field - first-author work at venues such as NeurIPS, ICML, ICLR or ACL, including the datasets and benchmarks tracks</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nCallosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous models and hardware. Our team is developing the science that makes that possible: a principled, evidence-based method for building agentic systems automatically at unprecedented scale.\n\nThe team tackles these problems on two fronts. We build the tools to analyse and evaluate agentic systems rigorously enough to say where and why they go wrong, and we use what we learn to design better ones. We are not simply building a better harness; the best harness will change with every new model, task and generation of silicon. We are building the layer beneath it, so that design decisions follow from evidence rather than taste.\n\nThis role owns the first of those fronts. Agentic evaluation today is not up to the job: benchmarks saturate, scores move for reasons unrelated to capability, results fail to reproduce across runs, and a single number tells you nothing about which step went wrong. Every design decision the rest of the team makes rests on the quality of that instrument. You will build measures whose construct we can defend, whose variance we understand, and which localise a failure rather than merely scoring it.\n\nThe mandate is broad, and most of the questions inside it are still unanswered. You will have wide latitude to choose problems, and your results will shape what the company builds in the future.\n\n\n\nWhat You'll Build\n\n - Design benchmarks and evaluation suites for agentic behaviour – multi-turn, long-horizon, tool-using, operating in non-stationary environments – including the properties that are hard to score, such as reliability and recovery from error\n\n - Build the methodology as well as the harness: statistical power, variance across runs, contamination and construct validity\n\n - Be the independent voice on the quality of our own results. Red-team our evaluations, design the sanity checks that catch a flattering number before it reaches a decision.\n\n - Turn raw traces into structured evidence. Failure taxonomies, behavioural signatures, and attribution of an outcome to the decision that caused it\n\n - Move from measurement to prediction: infer what a system can do from partial evidence, and estimate performance on a task before running it. This rests on defining what each model is genuinely good at, what benchmarks really measure, and what capability profile a real compound task demands\n\n - Build evaluation and observability infrastructure as durable instruments the whole company works from\n   \n   \n\nWhat You'll Bring\n\n - Evidence that you can run research of your own: you take an open question, design the experiments that settle it, and produce results other people can build on\n\n - Deep hands-on experience with LLMs in agentic settings: multi-step tasks, tool use, long horizons, and the specific ways all of it breaks\n\n - Statistical discipline. Hypotheses stated in advance, uncertainty quantified, and running comparisons that isolate one variable\n\n - Strong engineering. You write the code that runs your experiments and you are comfortable working inside a large shared codebase\n\n - Strong communication skills. You are able to turn research findings into clear, prioritised guidance for the teams who will act on them, and to write them up for a wider audience\n\nWhat Sets You Apart\n\n - Evaluations or benchmarks you built that other people went on to use\n\n - Broad training in empirical method, potentially from a field outside machine learning - experimental design, Bayesian inference, model selection, significance testing, uncertainty quantification\n\n - Experience with automated grading and model-based judging, and a rigorous account of its failure modes and its agreement with human raters\n\n - A published track record in a relevant field - first-author work at venues such as NeurIPS, ICML, ICLR or ACL, including the datasets and benchmarks tracks\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"8c29b53f-f173-4192-b61d-952c577bd001","title":"Evolutionary Optimisation - Member of Technical Staff","department":"Compute & Infrastructure","team":"Compute & Infrastructure","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:33:33.528+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/8c29b53f-f173-4192-b61d-952c577bd001","applyUrl":"https://jobs.ashbyhq.com/callosum/8c29b53f-f173-4192-b61d-952c577bd001/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that the next generation of AI systems will be co-optimised across models, software, and hardware. As we operate across increasingly heterogeneous workloads and accelerators, the number of possible system configurations grows combinatorially. Manually exploring that space does not scale – we need systems that can search it themselves.</p><p style=\"min-height:1.5em\">This role owns Callosum’s program evolution and optimisation loop: an LLM-guided search system that iteratively generates, evaluates, and improves programs against verifiable objectives. We already use this approach across kernels, inference runtimes, scheduling and scaling policies, KV caching strategies, and more. You will develop the loop both as a reusable internal optimisation system and as a research object in its own right, improving how it searches, measuring what makes it effective, and pushing its capabilities on increasingly difficult optimisation and reasoning problems.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Develop the core program evolution loop, improving how candidates are generated, evaluated, selected, and evolved over long optimisation runs</p></li><li><p style=\"min-height:1.5em\">Design search strategies that balance exploration, exploitation, diversity, and sample efficiency across vast and irregular optimisation spaces</p></li><li><p style=\"min-height:1.5em\">Build continuous benchmarks that measure the optimiser across internal workloads and standalone problems, making improvements and regressions empirically visible</p></li><li><p style=\"min-height:1.5em\">Turn the system into reliable internal infrastructure that can optimise new verifiable domains as a reusable internal tool that other team members can scale to new problems</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Strong foundations in optimisation and search, including familiarity with evolutionary, stochastic, combinatorial, or other non-gradient-based methods</p></li><li><p style=\"min-height:1.5em\">Ability to reason about search itself i.e. exploration versus exploitation, diversity, pruning, objective design, and when different optimisation strategies are appropriate</p></li><li><p style=\"min-height:1.5em\">Strong engineering instincts, with experience turning experimental systems into reusable, measurable, and reliable software</p></li><li><p style=\"min-height:1.5em\">A demonstrable tendency to draw from techniques across optimisation, algorithms, and machine learning rather than treating LLMs as the solution to every problem</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nCallosum believes that the next generation of AI systems will be co-optimised across models, software, and hardware. As we operate across increasingly heterogeneous workloads and accelerators, the number of possible system configurations grows combinatorially. Manually exploring that space does not scale – we need systems that can search it themselves.\n\nThis role owns Callosum’s program evolution and optimisation loop: an LLM-guided search system that iteratively generates, evaluates, and improves programs against verifiable objectives. We already use this approach across kernels, inference runtimes, scheduling and scaling policies, KV caching strategies, and more. You will develop the loop both as a reusable internal optimisation system and as a research object in its own right, improving how it searches, measuring what makes it effective, and pushing its capabilities on increasingly difficult optimisation and reasoning problems.\n\n\n\nWhat You'll Build\n\n - Develop the core program evolution loop, improving how candidates are generated, evaluated, selected, and evolved over long optimisation runs\n\n - Design search strategies that balance exploration, exploitation, diversity, and sample efficiency across vast and irregular optimisation spaces\n\n - Build continuous benchmarks that measure the optimiser across internal workloads and standalone problems, making improvements and regressions empirically visible\n\n - Turn the system into reliable internal infrastructure that can optimise new verifiable domains as a reusable internal tool that other team members can scale to new problems\n   \n   \n\nWhat Sets You Apart\n\n - Strong foundations in optimisation and search, including familiarity with evolutionary, stochastic, combinatorial, or other non-gradient-based methods\n\n - Ability to reason about search itself i.e. exploration versus exploitation, diversity, pruning, objective design, and when different optimisation strategies are appropriate\n\n - Strong engineering instincts, with experience turning experimental systems into reusable, measurable, and reliable software\n\n - A demonstrable tendency to draw from techniques across optimisation, algorithms, and machine learning rather than treating LLMs as the solution to every problem\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"71c6dcb8-4dda-4511-ab88-61523b8832fa","title":"Research Engineer, Benchmarking - Member of Technical Staff","department":"Applied AI","team":"Applied AI","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:33:00.003+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/71c6dcb8-4dda-4511-ab88-61523b8832fa","applyUrl":"https://jobs.ashbyhq.com/callosum/71c6dcb8-4dda-4511-ab88-61523b8832fa/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Choosing between algorithmic strategies for multi-step LLM work is a measurement problem, and most teams solve it badly: comparisons run case by case, by whoever needs them that week, on whatever task is closest to hand. That doesn't scale, and it doesn't hold up to outside scrutiny - from a customer, or from a reviewer. Callosum needs one benchmarking system: reproducible, contamination-controlled, and trusted enough to be the evidence that decides which approach ships.</p><p style=\"min-height:1.5em\">This role owns that system. You will build a harness that measures task success, quality, and robustness across motifs, agent topologies, and decomposition strategies, grounded in execution - real commits, real traces, sandboxed grading - rather than self-reported or model-graded scores. The results become the proof points we show customers, the evidence behind the benchmarks we co-publish, and the basis on which an approach ships or doesn't.</p><p style=\"min-height:1.5em\">This is a research hire that builds. We expect the rigour of a strong evaluation paper applied to a production system, and the engineering ability to design, build, and curate it yourself rather than hand it off.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Design and build a unified system for evaluating agentic and algorithmic solutions - task success, quality, and robustness across motifs, agent topologies, and decomposition strategies, on workloads that match what customers actually run. Cost per resolved task is an outcome you track, not the object of the exercise.</p></li><li><p style=\"min-height:1.5em\">Mine real commits and traces, run sandboxed execution grading, and build task suites that reflect real agentic work: code search, code edit and repair, repository summarisation, tool use. Self-reported or model-graded success isn't enough on its own.</p></li><li><p style=\"min-height:1.5em\">Enforce controls against contamination, overfitting to benchmarks, and metric gaming, and keep baselines stable over time - any result should be re-runnable to the same number, by us or by a reviewer</p></li><li><p style=\"min-height:1.5em\">Compare algorithmic and agentic approaches honestly, not models or chips - a motif that adds steps, latency, or cost has to earn it in resolved-task quality, and the system says clearly when it doesn't</p></li><li><p style=\"min-height:1.5em\">Lead external benchmark co-publications, held to a standard that survives peer and customer review</p></li><li><p style=\"min-height:1.5em\">Feed results directly into which approach ships, into the proof points behind customer engagements, and review quality claims across the company before they go out</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What You'll Bring</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">PhD in computer science, machine learning, or a related field, or an equivalent research track record</p></li><li><p style=\"min-height:1.5em\">Authorship or co-authorship of a benchmark or evaluation paper at a recognised venue - NeurIPS Datasets and Benchmarks, ICML, ICLR, ACL - ideally on agentic or LLM evaluation, or a comparably rigorous evaluation contribution</p></li><li><p style=\"min-height:1.5em\">A working understanding of how LLM and agent evaluation goes wrong: contamination, overfitting to benchmarks, weak baselines, underpowered comparisons, irreproducible results</p></li><li><p style=\"min-height:1.5em\">The engineering ability to design, build, and curate these systems decisively - strong Python, and comfort with sandboxed and distributed execution and CI</p></li><li><p style=\"min-height:1.5em\">Hands-on experience building or rigorously evaluating agentic or multi-step LLM systems</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Published agentic or tool-use benchmarks that use execution-based grading</p></li><li><p style=\"min-height:1.5em\">Experience running sandboxed execution grading at scale</p></li><li><p style=\"min-height:1.5em\">Open-source evaluation or harness tooling</p></li><li><p style=\"min-height:1.5em\">Familiarity with code-agent workloads such as search, edit, and repair</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nChoosing between algorithmic strategies for multi-step LLM work is a measurement problem, and most teams solve it badly: comparisons run case by case, by whoever needs them that week, on whatever task is closest to hand. That doesn't scale, and it doesn't hold up to outside scrutiny - from a customer, or from a reviewer. Callosum needs one benchmarking system: reproducible, contamination-controlled, and trusted enough to be the evidence that decides which approach ships.\n\nThis role owns that system. You will build a harness that measures task success, quality, and robustness across motifs, agent topologies, and decomposition strategies, grounded in execution - real commits, real traces, sandboxed grading - rather than self-reported or model-graded scores. The results become the proof points we show customers, the evidence behind the benchmarks we co-publish, and the basis on which an approach ships or doesn't.\n\nThis is a research hire that builds. We expect the rigour of a strong evaluation paper applied to a production system, and the engineering ability to design, build, and curate it yourself rather than hand it off.\n\n\n\nWhat You'll Build\n\n - Design and build a unified system for evaluating agentic and algorithmic solutions - task success, quality, and robustness across motifs, agent topologies, and decomposition strategies, on workloads that match what customers actually run. Cost per resolved task is an outcome you track, not the object of the exercise.\n\n - Mine real commits and traces, run sandboxed execution grading, and build task suites that reflect real agentic work: code search, code edit and repair, repository summarisation, tool use. Self-reported or model-graded success isn't enough on its own.\n\n - Enforce controls against contamination, overfitting to benchmarks, and metric gaming, and keep baselines stable over time - any result should be re-runnable to the same number, by us or by a reviewer\n\n - Compare algorithmic and agentic approaches honestly, not models or chips - a motif that adds steps, latency, or cost has to earn it in resolved-task quality, and the system says clearly when it doesn't\n\n - Lead external benchmark co-publications, held to a standard that survives peer and customer review\n\n - Feed results directly into which approach ships, into the proof points behind customer engagements, and review quality claims across the company before they go out\n   \n   \n\nWhat You'll Bring\n\n - PhD in computer science, machine learning, or a related field, or an equivalent research track record\n\n - Authorship or co-authorship of a benchmark or evaluation paper at a recognised venue - NeurIPS Datasets and Benchmarks, ICML, ICLR, ACL - ideally on agentic or LLM evaluation, or a comparably rigorous evaluation contribution\n\n - A working understanding of how LLM and agent evaluation goes wrong: contamination, overfitting to benchmarks, weak baselines, underpowered comparisons, irreproducible results\n\n - The engineering ability to design, build, and curate these systems decisively - strong Python, and comfort with sandboxed and distributed execution and CI\n\n - Hands-on experience building or rigorously evaluating agentic or multi-step LLM systems\n\n\n\nWhat Sets You Apart\n\n - Published agentic or tool-use benchmarks that use execution-based grading\n\n - Experience running sandboxed execution grading at scale\n\n - Open-source evaluation or harness tooling\n\n - Familiarity with code-agent workloads such as search, edit, and repair\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"ab687a5e-300e-419a-b1b5-449013680e81","title":"Security Lead - Member of Technical Staff","department":"Office of the CTO","team":"Office of the CTO","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:36:00.633+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/ab687a5e-300e-419a-b1b5-449013680e81","applyUrl":"https://jobs.ashbyhq.com/callosum/ab687a5e-300e-419a-b1b5-449013680e81/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum operates across a broad technical surface: APIs used by customers directly, orchestration software deployed across heterogeneous clusters, agentic systems with access to tools and data, and physical compute infrastructure that we operate ourselves. These products run in very different environments and for customers with very different security requirements, from shared cloud infrastructure through to sovereign, on-premise, and isolated deployments.</p><p style=\"min-height:1.5em\">This role owns how security is designed and assured across that landscape. You will set the security architecture, translate customer and deployment requirements into concrete technical expectations for engineering teams, and ensure that the guarantees we make can be demonstrated in practice. You will lead security engagements across customers, infrastructure partners, and external technical reviewers, while building and managing the team responsible for security across Callosum.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Do</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Define security architecture and requirements across our API, orchestration stack, agentic systems, and physical infrastructure</p></li><li><p style=\"min-height:1.5em\">Translate complex security requirements into specific engineering work, measurable guarantees, and verifiable evidence</p></li><li><p style=\"min-height:1.5em\">Lead security assurance for sensitive deployments from initial requirements through technical review and approval</p></li><li><p style=\"min-height:1.5em\">Build and lead the security team across architecture, threat modelling, adversarial evaluation, incident response, and operations</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience personally owning a demanding security assurance or sensitive deployment process end-to-end</p></li><li><p style=\"min-height:1.5em\">Strong technical foundations in infrastructure or security engineering, with the depth to challenge architectural decisions</p></li><li><p style=\"min-height:1.5em\">Experience securing sovereign, on-premise, air-gapped, multi-tenant, or other high-assurance environments</p></li><li><p style=\"min-height:1.5em\">Experience building and leading security teams, with strong judgement around risk, prioritisation, and engineering trade-offs</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nCallosum operates across a broad technical surface: APIs used by customers directly, orchestration software deployed across heterogeneous clusters, agentic systems with access to tools and data, and physical compute infrastructure that we operate ourselves. These products run in very different environments and for customers with very different security requirements, from shared cloud infrastructure through to sovereign, on-premise, and isolated deployments.\n\nThis role owns how security is designed and assured across that landscape. You will set the security architecture, translate customer and deployment requirements into concrete technical expectations for engineering teams, and ensure that the guarantees we make can be demonstrated in practice. You will lead security engagements across customers, infrastructure partners, and external technical reviewers, while building and managing the team responsible for security across Callosum.\n\n\n\nWhat You'll Do\n\n - Define security architecture and requirements across our API, orchestration stack, agentic systems, and physical infrastructure\n\n - Translate complex security requirements into specific engineering work, measurable guarantees, and verifiable evidence\n\n - Lead security assurance for sensitive deployments from initial requirements through technical review and approval\n\n - Build and lead the security team across architecture, threat modelling, adversarial evaluation, incident response, and operations\n   \n   \n\nWhat Sets You Apart\n\n - Experience personally owning a demanding security assurance or sensitive deployment process end-to-end\n\n - Strong technical foundations in infrastructure or security engineering, with the depth to challenge architectural decisions\n\n - Experience securing sovereign, on-premise, air-gapped, multi-tenant, or other high-assurance environments\n\n - Experience building and leading security teams, with strong judgement around risk, prioritisation, and engineering trade-offs\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."},{"id":"e386b6dd-5cb6-4f54-a690-87f136f9c8e7","title":"ML Research Engineer - Member of Technical Staff","department":"General Capabilities","team":"General Capabilities","employmentType":"FullTime","location":"London","secondaryLocations":[],"publishedAt":"2026-08-20T06:33:59.349+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"London","addressCountry":"United Kingdom","addressLocality":"London"}},"jobUrl":"https://jobs.ashbyhq.com/callosum/e386b6dd-5cb6-4f54-a690-87f136f9c8e7","applyUrl":"https://jobs.ashbyhq.com/callosum/e386b6dd-5cb6-4f54-a690-87f136f9c8e7/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>About Us</strong></p><p style=\"min-height:1.5em\">We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.</p><p style=\"min-height:1.5em\">Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.</p><p style=\"min-height:1.5em\">The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.</p><p style=\"min-height:1.5em\">Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.</p><p style=\"min-height:1.5em\">Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.</p><p style=\"min-height:1.5em\">In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.</p><p style=\"min-height:1.5em\">We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>About the Role</strong></p><p style=\"min-height:1.5em\">Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous models and hardware. Our team is developing the science that makes that possible: a principled, evidence-based method for building agentic systems automatically at unprecedented scale.</p><p style=\"min-height:1.5em\">Today, designing these systems is a craft. Decisions about memory management, task decomposition, tool use and inter-agent coordination are made on intuition and convention. The results work until they don't, and it is rarely clear in advance which failure comes next.</p><p style=\"min-height:1.5em\">The team tackles these problems on two fronts. We build the tools to analyse and evaluate agentic systems rigorously enough to say where and why they go wrong, and we use what we learn to design better ones. We are not simply building a better harness; the best harness will change with every new model, task and generation of silicon. We are building the layer beneath it, so that design decisions follow from evidence rather than taste.</p><p style=\"min-height:1.5em\">The mandate is broad, and most of the questions inside it are still unanswered. You will have wide latitude to choose problems, and your results will shape what the company builds in the future.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What You'll Build</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Design and run experiments that isolate why agentic systems fail - across behaviour, traces and activations - and turn those findings into interventions that measurably move intelligence, cost and reliability</p></li><li><p style=\"min-height:1.5em\">Attack one or more of our core research themes: context and memory management; agent steering, task decomposition and specialisation; continual learning during deployment; and inter-agent communication</p></li><li><p style=\"min-height:1.5em\">Build agent analysis and evaluation infrastructure - observability, behaviour analysis, evaluation harnesses, auditing - as durable instruments the whole company works from, not one-off scripts</p></li><li><p style=\"min-height:1.5em\">Publish. Evaluation results, methods and discoveries, as both company assets and public evidence</p></li><li><p style=\"min-height:1.5em\">Work across team boundaries in both directions: take a vertical domain-specific application problem and design systems which run at unpreceding scale at lowest cost and latency; or invent novel model and agent architecture which exploit emerging unconventional silicons beyond GPU</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What You'll Bring</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Evidence that you can run research of your own: you take an open question, design the experiments that settle it, and produce results other people can build on. Where you learned to do that matters less to us than that you can</p></li><li><p style=\"min-height:1.5em\">Deep hands-on experience with LLMs in agentic settings: multi-step tasks, tool use, long horizons, and the specific ways all of it breaks</p></li><li><p style=\"min-height:1.5em\">Real experimental discipline. Stated hypotheses, controlled comparisons, ablations that isolate one variable, honest uncertainty, and the instinct to distinguish a genuine effect from prompt luck or benchmark noise</p></li><li><p style=\"min-height:1.5em\">Strong engineering. You write the code that runs your own experiments and you are comfortable working inside a substantial shared codebase</p></li><li><p style=\"min-height:1.5em\">Strong communication skills. Able to turn research findings into clear, prioritised guidance for the teams who will act on them, and to write them up for a wider audience</p></li></ul><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>What Sets You Apart</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Evaluations, benchmarks or agent harnesses you built that other people went on to use</p></li><li><p style=\"min-height:1.5em\">Experience with RL or post-training for long-horizon, multi-step or tool-use tasks - reward design, environment construction, data generation</p></li><li><p style=\"min-height:1.5em\">Depth in multi-agent systems, planning, program synthesis, or retrieval over structured artefacts such as codebases</p></li><li><p style=\"min-height:1.5em\">Deep familiarity with the internals of SGLang, vLLM, or comparable inference serving frameworks - scheduler design, memory management, and execution pipelines</p></li><li><p style=\"min-height:1.5em\">A background in another field that studies systems of interacting heterogeneous components - neuroscience, distributed systems, economics, cognitive science - and the transfer of intuition that comes with it</p></li><li><p style=\"min-height:1.5em\">A published track record in a relevant field - first-author work at venues such as NeurIPS, ICML, ICLR , ACL or MLSys</p><p style=\"min-height:1.5em\"></p></li></ul><p style=\"min-height:1.5em\"><strong>What We Offer</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Competitive Salary, determined by skills and experience</p></li><li><p style=\"min-height:1.5em\">Equity &amp; Ownership</p></li><li><p style=\"min-height:1.5em\">Private healthcare</p></li><li><p style=\"min-height:1.5em\">We offer Visa sponsorship and relocation benefits to hire the best in the world</p></li><li><p style=\"min-height:1.5em\">We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us</p></li></ul><p style=\"min-height:1.5em\"><em>We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.</em></p>","descriptionPlain":"About Us\n\nWe’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.\n\nCallosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.\n\nThe last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.\n\nOur founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.\n\nBecause our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.\n\nIn our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.\n\nWe are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.\n\n\n\nAbout the Role\n\nCallosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous models and hardware. Our team is developing the science that makes that possible: a principled, evidence-based method for building agentic systems automatically at unprecedented scale.\n\nToday, designing these systems is a craft. Decisions about memory management, task decomposition, tool use and inter-agent coordination are made on intuition and convention. The results work until they don't, and it is rarely clear in advance which failure comes next.\n\nThe team tackles these problems on two fronts. We build the tools to analyse and evaluate agentic systems rigorously enough to say where and why they go wrong, and we use what we learn to design better ones. We are not simply building a better harness; the best harness will change with every new model, task and generation of silicon. We are building the layer beneath it, so that design decisions follow from evidence rather than taste.\n\nThe mandate is broad, and most of the questions inside it are still unanswered. You will have wide latitude to choose problems, and your results will shape what the company builds in the future.\n\n\n\nWhat You'll Build\n\n - Design and run experiments that isolate why agentic systems fail - across behaviour, traces and activations - and turn those findings into interventions that measurably move intelligence, cost and reliability\n\n - Attack one or more of our core research themes: context and memory management; agent steering, task decomposition and specialisation; continual learning during deployment; and inter-agent communication\n\n - Build agent analysis and evaluation infrastructure - observability, behaviour analysis, evaluation harnesses, auditing - as durable instruments the whole company works from, not one-off scripts\n\n - Publish. Evaluation results, methods and discoveries, as both company assets and public evidence\n\n - Work across team boundaries in both directions: take a vertical domain-specific application problem and design systems which run at unpreceding scale at lowest cost and latency; or invent novel model and agent architecture which exploit emerging unconventional silicons beyond GPU\n   \n   \n\nWhat You'll Bring\n\n - Evidence that you can run research of your own: you take an open question, design the experiments that settle it, and produce results other people can build on. Where you learned to do that matters less to us than that you can\n\n - Deep hands-on experience with LLMs in agentic settings: multi-step tasks, tool use, long horizons, and the specific ways all of it breaks\n\n - Real experimental discipline. Stated hypotheses, controlled comparisons, ablations that isolate one variable, honest uncertainty, and the instinct to distinguish a genuine effect from prompt luck or benchmark noise\n\n - Strong engineering. You write the code that runs your own experiments and you are comfortable working inside a substantial shared codebase\n\n - Strong communication skills. Able to turn research findings into clear, prioritised guidance for the teams who will act on them, and to write them up for a wider audience\n\n\n\nWhat Sets You Apart\n\n - Evaluations, benchmarks or agent harnesses you built that other people went on to use\n\n - Experience with RL or post-training for long-horizon, multi-step or tool-use tasks - reward design, environment construction, data generation\n\n - Depth in multi-agent systems, planning, program synthesis, or retrieval over structured artefacts such as codebases\n\n - Deep familiarity with the internals of SGLang, vLLM, or comparable inference serving frameworks - scheduler design, memory management, and execution pipelines\n\n - A background in another field that studies systems of interacting heterogeneous components - neuroscience, distributed systems, economics, cognitive science - and the transfer of intuition that comes with it\n\n - A published track record in a relevant field - first-author work at venues such as NeurIPS, ICML, ICLR , ACL or MLSys\n   \n   \n\nWhat We Offer\n\n - Competitive Salary, determined by skills and experience\n\n - Equity & Ownership\n\n - Private healthcare\n\n - We offer Visa sponsorship and relocation benefits to hire the best in the world\n\n - We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us\n\nWe're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all."}],"apiVersion":"1"}