{"jobs":[{"id":"c375554c-c33c-4599-9dc6-078ef2ae71e1","title":"Founding Engineer, Open Source","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"New York Office","secondaryLocations":[],"publishedAt":"2026-07-23T15:21:26.520+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"New York","addressCountry":"United States","addressLocality":"New York"}},"jobUrl":"https://jobs.ashbyhq.com/datalab/c375554c-c33c-4599-9dc6-078ef2ae71e1","applyUrl":"https://jobs.ashbyhq.com/datalab/c375554c-c33c-4599-9dc6-078ef2ae71e1/application","descriptionHtml":"<h1>Founding Engineer, Open Source</h1><p style=\"min-height:1.5em\"><strong>Salary range</strong> — $225k – $300k | <strong>Equity</strong> — 0.15-0.25% | <strong>In-person NYC</strong></p><p style=\"min-height:1.5em\"></p><h2>About Datalab</h2><p style=\"min-height:1.5em\">Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.</p><p style=\"min-height:1.5em\">We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, chandra, surya, marker, and lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.</p><p style=\"min-height:1.5em\"></p><h2>Role Overview</h2><p style=\"min-height:1.5em\">We're looking for a Founding Engineer to develop and evangelize our open source repos. This includes chandra, surya, marker, pdftext, and lift, which collectively have over 70k Github stars. It also includes new tools we have yet to build and launch.</p><p style=\"min-height:1.5em\">As you work on our repos, you’ll also become the credible person to evangelize them. You’ll turn your own work into demos, benchmarks, tutorials, and launches. You’ll also support other launches across Datalab, especially when they touch open source components like the SDK.</p><p style=\"min-height:1.5em\">Our projects have real reach: 70k+ GitHub stars and users everywhere from frontier AI labs to Fortune 500s. Your job is to turn that reach into a thriving, engaged developer community through content, code, and showing up where developers already are.</p><p style=\"min-height:1.5em\">If you're the kind of engineer who’s energized by both building things and helping other people build, this is the role for you.</p><p style=\"min-height:1.5em\"><strong>Day to day:</strong></p><p style=\"min-height:1.5em\"><strong>Own our open source repos and SDK:</strong> Drive chandra, surya, marker, lift, and our Python SDK forward as a hands-on contributor. Ship new features and improvements that the market cares about, help plan new versions, and own open-source launches end-to-end. This spans from the release itself to the demos, benchmarks, and content to evangelize the launch.</p><p style=\"min-height:1.5em\"><strong>Build demos and benchmarks.</strong> Build example apps, templates, and starter projects on top of Datalab's models and API that make it obvious what's possible and easy to remix. Publish benchmarks and evals that show, in code, how our models compare.</p><p style=\"min-height:1.5em\"><strong>Turn your work into content that teaches:</strong> Write technical blogs, guides, tutorials, and changelogs that help developers get started with marker, surya, chandra, lift, and our API, then go deep. Establish cookbooks and quickstart guides where we don't have them yet. Write content that helps developers differentiate between our platform / our models and our competitors (e.g., how to run evals).</p><p style=\"min-height:1.5em\"><strong>Close the loop with product:</strong> Bring developer feedback, friction points, and emerging needs back to engineering and product to help shape what we build next.</p><p style=\"min-height:1.5em\"><strong>Grow the community:</strong> Engage developers on GitHub, Discord, X, Reddit, and beyond. Celebrate contributors, answer questions, highlight cool projects, and add value in conversations rather than just promote. Be a visible, trusted voice for the platform online.</p><p style=\"min-height:1.5em\"></p><h2>Ideal Candidate</h2><p style=\"min-height:1.5em\">You're a strong engineer who’s shipped things developers use, and you also love helping other people build. You know how to explain a complex idea simply, and you're as comfortable writing a tutorial or recording a demo as you are building a feature.</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">5+ years building production software, with meaningful open-source contributions</p></li><li><p style=\"min-height:1.5em\">Strong hands-on engineering skills; fluent in Python, shipping real, reproducible, well-tested code</p></li><li><p style=\"min-height:1.5em\">Track record of owning code end-to-end - design, implementation, release, and maintenance</p></li><li><p style=\"min-height:1.5em\">Able to create credible technical content - blogs, docs, tutorials, demos, or video - that developers actually read, watch, and use, because you understand the system deeply</p></li><li><p style=\"min-height:1.5em\">Comfortable representing your work publicly, whether in writing, on camera, or in online communities</p></li><li><p style=\"min-height:1.5em\">Comfortable with an early-stage startup - self-directed, hands-on, and happy to balance depth with shipping velocity</p></li></ul><p style=\"min-height:1.5em\"><strong>Bonus points if you:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Have experience with OCR, document AI, or structured extraction</p></li><li><p style=\"min-height:1.5em\">Have contributed to or maintained a widely-used open-source ML/vision/NLP project</p></li><li><p style=\"min-height:1.5em\">Have a growing audience on YouTube, X, or Twitch, or an active presence in developer communities</p></li><li><p style=\"min-height:1.5em\">Have used (and loved) marker, surya, chandra, or lift</p></li></ul><h2>Interview process</h2><ol style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">30-minute video call to evaluate fit</p></li><li><p style=\"min-height:1.5em\">1.5-hour in person meeting to build an architecture together</p></li><li><p style=\"min-height:1.5em\">Async code sharing + review phase (not a take-home project, you just share some code samples)</p></li><li><p style=\"min-height:1.5em\">Culture fit interviews with the team</p></li></ol><p style=\"min-height:1.5em\">At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.</p><p style=\"min-height:1.5em\"></p><h1>Equal opportunity</h1><p style=\"min-height:1.5em\">Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.</p><p style=\"min-height:1.5em\">If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.</p>","descriptionPlain":"FOUNDING ENGINEER, OPEN SOURCE\n\nSalary range — $225k – $300k | Equity — 0.15-0.25% | In-person NYC\n\n\n\n\nABOUT DATALAB\n\nDatalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.\n\nWe’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, chandra, surya, marker, and lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.\n\n\n\n\nROLE OVERVIEW\n\nWe're looking for a Founding Engineer to develop and evangelize our open source repos. This includes chandra, surya, marker, pdftext, and lift, which collectively have over 70k Github stars. It also includes new tools we have yet to build and launch.\n\nAs you work on our repos, you’ll also become the credible person to evangelize them. You’ll turn your own work into demos, benchmarks, tutorials, and launches. You’ll also support other launches across Datalab, especially when they touch open source components like the SDK.\n\nOur projects have real reach: 70k+ GitHub stars and users everywhere from frontier AI labs to Fortune 500s. Your job is to turn that reach into a thriving, engaged developer community through content, code, and showing up where developers already are.\n\nIf you're the kind of engineer who’s energized by both building things and helping other people build, this is the role for you.\n\nDay to day:\n\nOwn our open source repos and SDK: Drive chandra, surya, marker, lift, and our Python SDK forward as a hands-on contributor. Ship new features and improvements that the market cares about, help plan new versions, and own open-source launches end-to-end. This spans from the release itself to the demos, benchmarks, and content to evangelize the launch.\n\nBuild demos and benchmarks. Build example apps, templates, and starter projects on top of Datalab's models and API that make it obvious what's possible and easy to remix. Publish benchmarks and evals that show, in code, how our models compare.\n\nTurn your work into content that teaches: Write technical blogs, guides, tutorials, and changelogs that help developers get started with marker, surya, chandra, lift, and our API, then go deep. Establish cookbooks and quickstart guides where we don't have them yet. Write content that helps developers differentiate between our platform / our models and our competitors (e.g., how to run evals).\n\nClose the loop with product: Bring developer feedback, friction points, and emerging needs back to engineering and product to help shape what we build next.\n\nGrow the community: Engage developers on GitHub, Discord, X, Reddit, and beyond. Celebrate contributors, answer questions, highlight cool projects, and add value in conversations rather than just promote. Be a visible, trusted voice for the platform online.\n\n\n\n\nIDEAL CANDIDATE\n\nYou're a strong engineer who’s shipped things developers use, and you also love helping other people build. You know how to explain a complex idea simply, and you're as comfortable writing a tutorial or recording a demo as you are building a feature.\n\n - 5+ years building production software, with meaningful open-source contributions\n\n - Strong hands-on engineering skills; fluent in Python, shipping real, reproducible, well-tested code\n\n - Track record of owning code end-to-end - design, implementation, release, and maintenance\n\n - Able to create credible technical content - blogs, docs, tutorials, demos, or video - that developers actually read, watch, and use, because you understand the system deeply\n\n - Comfortable representing your work publicly, whether in writing, on camera, or in online communities\n\n - Comfortable with an early-stage startup - self-directed, hands-on, and happy to balance depth with shipping velocity\n\nBonus points if you:\n\n - Have experience with OCR, document AI, or structured extraction\n\n - Have contributed to or maintained a widely-used open-source ML/vision/NLP project\n\n - Have a growing audience on YouTube, X, or Twitch, or an active presence in developer communities\n\n - Have used (and loved) marker, surya, chandra, or lift\n\n\nINTERVIEW PROCESS\n\n 1. 30-minute video call to evaluate fit\n\n 2. 1.5-hour in person meeting to build an architecture together\n\n 3. Async code sharing + review phase (not a take-home project, you just share some code samples)\n\n 4. Culture fit interviews with the team\n\nAt this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.\n\n\n\n\nEQUAL OPPORTUNITY\n\nDatalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.\n\nIf you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to."},{"id":"fd783b8f-ea1d-4213-8f27-5f54d86dac10","title":"Senior Software Engineer","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"New York Office","secondaryLocations":[],"publishedAt":"2026-07-06T19:43:22.541+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"New York","addressCountry":"United States","addressLocality":"New York"}},"jobUrl":"https://jobs.ashbyhq.com/datalab/fd783b8f-ea1d-4213-8f27-5f54d86dac10","applyUrl":"https://jobs.ashbyhq.com/datalab/fd783b8f-ea1d-4213-8f27-5f54d86dac10/application","descriptionHtml":"<p style=\"min-height:1.5em\">Salary range: $225k - $300k | Equity: 0.15% - 0.35% | In-Person: NYC</p><p style=\"min-height:1.5em\"></p><h2><strong>About Datalab</strong></h2><p style=\"min-height:1.5em\">Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.</p><p style=\"min-height:1.5em\">We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><h2>Role Overview</h2><p style=\"min-height:1.5em\">We’re looking for a fullstack engineer who wants to build the interfaces, tools, and infrastructure that help developers and enterprises use our models. You’ll work across the stack to shape how people interact with OCR, extraction, and document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks, and improving the developer experience across our open-source repos and API.</p><p style=\"min-height:1.5em\">This is a high-ownership role that blends engineering, product thinking, and community engagement. You will work closely with the founders and the rest of the team to ship features, improve performance, and make our technology accessible to a global community of builders.</p><p style=\"min-height:1.5em\">As a small and fast-moving team, roles are fluid. You should enjoy working across backend, frontend, performance, and user-facing surfaces. Your work will directly influence how teams evaluate and deploy our models.</p><p style=\"min-height:1.5em\"></p><h2><strong>Day to day, you will:</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Ship features to our open source repos, API, and internal tooling.</p></li><li><p style=\"min-height:1.5em\">Design and build frontend features that make document parsing more interactive and understandable.</p></li><li><p style=\"min-height:1.5em\">Optimize inference performance and improve the reliability of our frameworks.</p></li><li><p style=\"min-height:1.5em\">Engage with the community on Github and Discord by collecting feedback, debugging issues, and identifying opportunities for improvement.</p></li><li><p style=\"min-height:1.5em\">Support customer implementations with technical guidance and troubleshooting when needed.</p></li></ul><p style=\"min-height:1.5em\"></p><h2>Ideal Candidate</h2><p style=\"min-height:1.5em\">You thrive at the intersection of engineering, product, and user experience. You like working close to real users and making technical systems feel simple and intuitive. You operate with autonomy and ownership and enjoy moving quickly in an environment where you can directly influence outcomes.</p><p style=\"min-height:1.5em\"><strong>We’re eager to work with someone who has:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">5+ years of fullstack development experience building APIs and/or developer-focused products.</p></li><li><p style=\"min-height:1.5em\">Experience building in early-stage startup environments.</p></li><li><p style=\"min-height:1.5em\">Shipped and maintained production systems serving high-volume traffic.</p></li><li><p style=\"min-height:1.5em\">Experience building with LLMs and agentic workflows.</p></li></ul><p style=\"min-height:1.5em\"><strong>Bonus points:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience with Python ML frameworks (PyTorch, Transformers, etc.) or interest in training models.</p></li><li><p style=\"min-height:1.5em\">Familiarity with document processing, computer vision, or OCR systems.</p></li><li><p style=\"min-height:1.5em\">Maintained or contributed to open-source projects with active communities</p></li><li><p style=\"min-height:1.5em\">Experience writing technical content or demos to showcase your work.</p></li></ul><p style=\"min-height:1.5em\"></p><h1>Interview process</h1><ol style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">30-minute video call to evaluate fit</p></li><li><p style=\"min-height:1.5em\">90-minute live architecture discussion</p></li><li><p style=\"min-height:1.5em\">Culture fit interviews with the team</p></li></ol><p style=\"min-height:1.5em\">At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.</p><p style=\"min-height:1.5em\"></p><h1>Equal opportunity</h1><p style=\"min-height:1.5em\">Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.</p><p style=\"min-height:1.5em\">If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.</p>","descriptionPlain":"Salary range: $225k - $300k | Equity: 0.15% - 0.35% | In-Person: NYC\n\n\n\n\nABOUT DATALAB\n\nDatalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.\n\nWe’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.\n\n \n\n\nROLE OVERVIEW\n\nWe’re looking for a fullstack engineer who wants to build the interfaces, tools, and infrastructure that help developers and enterprises use our models. You’ll work across the stack to shape how people interact with OCR, extraction, and document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks, and improving the developer experience across our open-source repos and API.\n\nThis is a high-ownership role that blends engineering, product thinking, and community engagement. You will work closely with the founders and the rest of the team to ship features, improve performance, and make our technology accessible to a global community of builders.\n\nAs a small and fast-moving team, roles are fluid. You should enjoy working across backend, frontend, performance, and user-facing surfaces. Your work will directly influence how teams evaluate and deploy our models.\n\n\n\n\nDAY TO DAY, YOU WILL:\n\n - Ship features to our open source repos, API, and internal tooling.\n\n - Design and build frontend features that make document parsing more interactive and understandable.\n\n - Optimize inference performance and improve the reliability of our frameworks.\n\n - Engage with the community on Github and Discord by collecting feedback, debugging issues, and identifying opportunities for improvement.\n\n - Support customer implementations with technical guidance and troubleshooting when needed.\n\n\n\n\nIDEAL CANDIDATE\n\nYou thrive at the intersection of engineering, product, and user experience. You like working close to real users and making technical systems feel simple and intuitive. You operate with autonomy and ownership and enjoy moving quickly in an environment where you can directly influence outcomes.\n\nWe’re eager to work with someone who has:\n\n - 5+ years of fullstack development experience building APIs and/or developer-focused products.\n\n - Experience building in early-stage startup environments.\n\n - Shipped and maintained production systems serving high-volume traffic.\n\n - Experience building with LLMs and agentic workflows.\n\nBonus points:\n\n - Experience with Python ML frameworks (PyTorch, Transformers, etc.) or interest in training models.\n\n - Familiarity with document processing, computer vision, or OCR systems.\n\n - Maintained or contributed to open-source projects with active communities\n\n - Experience writing technical content or demos to showcase your work.\n\n\n\n\nINTERVIEW PROCESS\n\n 1. 30-minute video call to evaluate fit\n\n 2. 90-minute live architecture discussion\n\n 3. Culture fit interviews with the team\n\nAt this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.\n\n\n\n\nEQUAL OPPORTUNITY\n\nDatalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.\n\nIf you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to."},{"id":"a974589b-6ffa-484a-84e2-7f527aef3293","title":"Founding GTM","department":"Sales","team":"Sales","employmentType":"FullTime","location":"New York Office","secondaryLocations":[],"publishedAt":"2026-07-06T20:04:24.398+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"New York","addressCountry":"United States","addressLocality":"New York"}},"jobUrl":"https://jobs.ashbyhq.com/datalab/a974589b-6ffa-484a-84e2-7f527aef3293","applyUrl":"https://jobs.ashbyhq.com/datalab/a974589b-6ffa-484a-84e2-7f527aef3293/application","descriptionHtml":"<h2><strong>About Datalab</strong></h2><p style=\"min-height:1.5em\">Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.</p><p style=\"min-height:1.5em\">We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><h2>Role Overview</h2><p style=\"min-height:1.5em\">We're hiring our founding GTM - someone who can run the full cycle of sales; sourcing leads, managing the sales process, and closing deals, all while building the playbook that future hires will run on.</p><p style=\"min-height:1.5em\">Datalab makes document AI infrastructure that powers extraction at scale. We're at 8-figure revenue, and have grown revenue &gt;5x YoY, with a team of 7. Anthropic uses Datalab. So do hundreds of other companies across FAANG, frontier AI labs, financial services, insurance, logistics, healthcare, and government. Our open-source projects (Marker, Surya, Chandra) have 60k+ stars and wide community adoption.</p><p style=\"min-height:1.5em\">Sales today is founder-led. The goal of this role is to build a real sales motion on top of that foundation - outbound, ICP definition, enterprise process, and the playbook itself. You won’t be selling alone - engineers and the founder are heavily involved in the sales process, and the whole company pitches in to unblock deals and support customers.</p><p style=\"min-height:1.5em\">This is a high-ownership, high-ambiguity role. We have standard pricing in some areas and open questions in others. We have strong signals about who our best customers are, but ICP isn't fully defined. You'll work alongside the founder and our GTM team to figure those things out by closing deals.</p><p style=\"min-height:1.5em\">If you want to inherit a polished playbook, this isn't the right role. If you want to build that playbook at a company with real product-market fit, a steady inbound pipeline, and serious enterprise customers, this is a great fit. There's a clear path to sales leadership for the person who builds the function well.</p><p style=\"min-height:1.5em\"></p><h2>Day to day you will:</h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Run the inbound deal cycle at volume on enterprise opportunities - discovery, demos, technical evaluation, pricing, security review, procurement, and close.</p></li><li><p style=\"min-height:1.5em\">Own the customer through successful deployment - you're the primary point of contact through onboarding and go-live. After deployment, accounts move to team support for reactive issues; you retain expansion and renewal on accounts you closed.</p></li><li><p style=\"min-height:1.5em\">Prioritize ruthlessly for enterprise accounts. Disqualify smaller inbound back to the self-serve flow, so we can learn from it.</p></li><li><p style=\"min-height:1.5em\">Define our ICP through every closed-won and closed-lost, in close partnership with the founder and chief of staff.</p></li><li><p style=\"min-height:1.5em\">Navigate technical sales - know when to loop in engineering for a POC, when to bring in the founder, and how to drive procurement, legal, and public-sector vehicles like GSA/SEWP.</p></li><li><p style=\"min-height:1.5em\">Keep CRM clean - every deal, every stage, every next step, forecasts that reflect reality.</p></li><li><p style=\"min-height:1.5em\">Build the outbound motion as inbound capacity stabilizes - identify high-fit accounts, run sequences, measure conversion, and double down on what works. By month 3 this should be producing meaningful self-sourced pipeline.</p></li><li><p style=\"min-height:1.5em\">Build the playbook - discovery scripts, demo flows, objection libraries, competitive battle cards, pricing guidance, qualification criteria.</p></li><li><p style=\"min-height:1.5em\">Close the product feedback loop - track what customers are blocked on and which features would unblock deals, prioritize by revenue impact, and bring structured signal to product and engineering on a regular cadence.</p></li></ul><p style=\"min-height:1.5em\"></p><h2>What success looks like</h2><h3>By month 3</h3><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Running the inbound deal cycle independently for most enterprise accounts.</p></li><li><p style=\"min-height:1.5em\">Cycle time on inbound deals shorter than today's baseline.</p></li><li><p style=\"min-height:1.5em\">Outbound experiments live, with first self-sourced pipeline starting to land.</p></li></ul><h3>By month 12</h3><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Hitting ramped quota.</p></li><li><p style=\"min-height:1.5em\">Outbound channel producing meaningful pipeline, with first self-sourced closes on the board.</p></li><li><p style=\"min-height:1.5em\">Playbook v1 documented, named-account list and outbound sequences in a state the next hire can inherit.</p></li></ul><p style=\"min-height:1.5em\"></p><h2>Ideal candidate</h2><p style=\"min-height:1.5em\">You think of yourself as a builder who happens to be in sales. You're energized by the fact that nothing is fully figured out yet, and you'd rather invent the system than inherit one. You take satisfaction in process - clean CRM hygiene, weekly metrics reviews, postmortems on lost deals.</p><p style=\"min-height:1.5em\">You're technical enough to hold your own. Our buyers are engineers, ML leads, and CTOs. You should be able to ramp fast on what Marker, Surya, and Chandra actually do, talk credibly about benchmarks and evals, and know when you're at the edge of your depth and need to bring in an engineer.</p><p style=\"min-height:1.5em\">You're hungry. You read your own call recordings. You ask the engineering team how the model handles a specific edge case the prospect raised. You come back to the next call with a sharper answer.</p><p style=\"min-height:1.5em\">You’re collaborative. You make the call that's right for the company and team, not just your commission - working a small-commission strategic logo because the case study compounds, pushing sub-fit inbound back to self-serve instead of grinding it for personal comp, flagging when the brand or engineering team did the heavy lifting on a deal. You trust that doing the right thing grows the pie for everyone - including yourself.</p><p style=\"min-height:1.5em\"></p><h2>We're eager to work with someone who has:</h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">5+ years in B2B SaaS sales with meaningful time in mid-market or enterprise.</p></li><li><p style=\"min-height:1.5em\">Owned complex, multi-stakeholder deals end-to-end - technical evaluation, security review, legal, procurement.</p></li><li><p style=\"min-height:1.5em\">Built outbound from scratch or rebuilt a broken motion, with clear metrics on what worked.</p></li><li><p style=\"min-height:1.5em\">Sold a technical product (APIs, infra, ML/AI tooling, developer platforms) to technical buyers.</p></li><li><p style=\"min-height:1.5em\">Operated in early-stage environments where the playbook didn't exist yet.</p></li><li><p style=\"min-height:1.5em\">Track record of meticulous CRM hygiene and rigorous pipeline math.</p></li><li><p style=\"min-height:1.5em\">Can use AI tools like Claude and Cowork effectively to streamline your work and build playbooks.</p><p style=\"min-height:1.5em\"></p></li></ul><h2>Bonus points if you:</h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Have sold document AI, OCR, IDP, or data extraction.</p></li><li><p style=\"min-height:1.5em\">Have closed deals in regulated verticals and are familiar with SOC 2, HIPAA, FedRAMP, DPAs, BAAs.</p></li><li><p style=\"min-height:1.5em\">Have sold both API/usage-based products and on-prem deployments.</p></li><li><p style=\"min-height:1.5em\">Have a technical background (CS degree, engineering experience, or strong self-taught fluency).</p></li><li><p style=\"min-height:1.5em\">Actively use open-source LLMs and/or projects.</p></li></ul><p style=\"min-height:1.5em\"></p><h2>Interview process</h2><ol style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">30-minute intro call with the founder.</p></li><li><p style=\"min-height:1.5em\">Sales process deep-dive - walk us through a complex deal you've owned end-to-end.</p></li><li><p style=\"min-height:1.5em\">Mock discovery + demo using our product, with discussion of a 1-page outbound strategy you'll submit beforehand (we'll provide the brief).</p></li><li><p style=\"min-height:1.5em\">Onsite half-day in NYC with the team.</p></li></ol><p style=\"min-height:1.5em\">At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.</p><p style=\"min-height:1.5em\"></p><h1>Equal opportunity</h1><p style=\"min-height:1.5em\">Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.</p><p style=\"min-height:1.5em\">If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.</p>","descriptionPlain":"ABOUT DATALAB\n\nDatalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.\n\nWe’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.\n\n \n\n\nROLE OVERVIEW\n\nWe're hiring our founding GTM - someone who can run the full cycle of sales; sourcing leads, managing the sales process, and closing deals, all while building the playbook that future hires will run on.\n\nDatalab makes document AI infrastructure that powers extraction at scale. We're at 8-figure revenue, and have grown revenue >5x YoY, with a team of 7. Anthropic uses Datalab. So do hundreds of other companies across FAANG, frontier AI labs, financial services, insurance, logistics, healthcare, and government. Our open-source projects (Marker, Surya, Chandra) have 60k+ stars and wide community adoption.\n\nSales today is founder-led. The goal of this role is to build a real sales motion on top of that foundation - outbound, ICP definition, enterprise process, and the playbook itself. You won’t be selling alone - engineers and the founder are heavily involved in the sales process, and the whole company pitches in to unblock deals and support customers.\n\nThis is a high-ownership, high-ambiguity role. We have standard pricing in some areas and open questions in others. We have strong signals about who our best customers are, but ICP isn't fully defined. You'll work alongside the founder and our GTM team to figure those things out by closing deals.\n\nIf you want to inherit a polished playbook, this isn't the right role. If you want to build that playbook at a company with real product-market fit, a steady inbound pipeline, and serious enterprise customers, this is a great fit. There's a clear path to sales leadership for the person who builds the function well.\n\n\n\n\nDAY TO DAY YOU WILL:\n\n - Run the inbound deal cycle at volume on enterprise opportunities - discovery, demos, technical evaluation, pricing, security review, procurement, and close.\n\n - Own the customer through successful deployment - you're the primary point of contact through onboarding and go-live. After deployment, accounts move to team support for reactive issues; you retain expansion and renewal on accounts you closed.\n\n - Prioritize ruthlessly for enterprise accounts. Disqualify smaller inbound back to the self-serve flow, so we can learn from it.\n\n - Define our ICP through every closed-won and closed-lost, in close partnership with the founder and chief of staff.\n\n - Navigate technical sales - know when to loop in engineering for a POC, when to bring in the founder, and how to drive procurement, legal, and public-sector vehicles like GSA/SEWP.\n\n - Keep CRM clean - every deal, every stage, every next step, forecasts that reflect reality.\n\n - Build the outbound motion as inbound capacity stabilizes - identify high-fit accounts, run sequences, measure conversion, and double down on what works. By month 3 this should be producing meaningful self-sourced pipeline.\n\n - Build the playbook - discovery scripts, demo flows, objection libraries, competitive battle cards, pricing guidance, qualification criteria.\n\n - Close the product feedback loop - track what customers are blocked on and which features would unblock deals, prioritize by revenue impact, and bring structured signal to product and engineering on a regular cadence.\n\n\n\n\nWHAT SUCCESS LOOKS LIKE\n\n\nBY MONTH 3\n\n - Running the inbound deal cycle independently for most enterprise accounts.\n\n - Cycle time on inbound deals shorter than today's baseline.\n\n - Outbound experiments live, with first self-sourced pipeline starting to land.\n\n\nBY MONTH 12\n\n - Hitting ramped quota.\n\n - Outbound channel producing meaningful pipeline, with first self-sourced closes on the board.\n\n - Playbook v1 documented, named-account list and outbound sequences in a state the next hire can inherit.\n\n\n\n\nIDEAL CANDIDATE\n\nYou think of yourself as a builder who happens to be in sales. You're energized by the fact that nothing is fully figured out yet, and you'd rather invent the system than inherit one. You take satisfaction in process - clean CRM hygiene, weekly metrics reviews, postmortems on lost deals.\n\nYou're technical enough to hold your own. Our buyers are engineers, ML leads, and CTOs. You should be able to ramp fast on what Marker, Surya, and Chandra actually do, talk credibly about benchmarks and evals, and know when you're at the edge of your depth and need to bring in an engineer.\n\nYou're hungry. You read your own call recordings. You ask the engineering team how the model handles a specific edge case the prospect raised. You come back to the next call with a sharper answer.\n\nYou’re collaborative. You make the call that's right for the company and team, not just your commission - working a small-commission strategic logo because the case study compounds, pushing sub-fit inbound back to self-serve instead of grinding it for personal comp, flagging when the brand or engineering team did the heavy lifting on a deal. You trust that doing the right thing grows the pie for everyone - including yourself.\n\n\n\n\nWE'RE EAGER TO WORK WITH SOMEONE WHO HAS:\n\n - 5+ years in B2B SaaS sales with meaningful time in mid-market or enterprise.\n\n - Owned complex, multi-stakeholder deals end-to-end - technical evaluation, security review, legal, procurement.\n\n - Built outbound from scratch or rebuilt a broken motion, with clear metrics on what worked.\n\n - Sold a technical product (APIs, infra, ML/AI tooling, developer platforms) to technical buyers.\n\n - Operated in early-stage environments where the playbook didn't exist yet.\n\n - Track record of meticulous CRM hygiene and rigorous pipeline math.\n\n - Can use AI tools like Claude and Cowork effectively to streamline your work and build playbooks.\n   \n   \n\n\nBONUS POINTS IF YOU:\n\n - Have sold document AI, OCR, IDP, or data extraction.\n\n - Have closed deals in regulated verticals and are familiar with SOC 2, HIPAA, FedRAMP, DPAs, BAAs.\n\n - Have sold both API/usage-based products and on-prem deployments.\n\n - Have a technical background (CS degree, engineering experience, or strong self-taught fluency).\n\n - Actively use open-source LLMs and/or projects.\n\n\n\n\nINTERVIEW PROCESS\n\n 1. 30-minute intro call with the founder.\n\n 2. Sales process deep-dive - walk us through a complex deal you've owned end-to-end.\n\n 3. Mock discovery + demo using our product, with discussion of a 1-page outbound strategy you'll submit beforehand (we'll provide the brief).\n\n 4. Onsite half-day in NYC with the team.\n\nAt this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.\n\n\n\n\nEQUAL OPPORTUNITY\n\nDatalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.\n\nIf you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to."},{"id":"7fac4a34-86d6-47cd-b44c-12ac0a7cb174","title":"Research Engineer","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"New York Office","secondaryLocations":[],"publishedAt":"2026-07-06T18:40:31.390+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"New York","addressCountry":"United States","addressLocality":"New York"}},"jobUrl":"https://jobs.ashbyhq.com/datalab/7fac4a34-86d6-47cd-b44c-12ac0a7cb174","applyUrl":"https://jobs.ashbyhq.com/datalab/7fac4a34-86d6-47cd-b44c-12ac0a7cb174/application","descriptionHtml":"<p style=\"min-height:1.5em\">Salary range - $275k - $325k | Equity - 0.25% | In-person NYC</p><p style=\"min-height:1.5em\"></p><h2>About Datalab</h2><p style=\"min-height:1.5em\">Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.</p><p style=\"min-height:1.5em\">We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.</p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><h2><strong>Role Overview</strong></h2><p style=\"min-height:1.5em\">We're looking for a Research Engineer to own problems end to end across our models, inference service, and product. You won't just train a model and hand it off. You'll take it from training through benchmarking, into our inference stack, and work with the team to integrate it into our products.</p><p style=\"min-height:1.5em\">We're a small team that has shipped the current state of the art OCR model, Chandra. Our models collectively have 70k+ Github stars. Our tools are used internally at frontier AI labs like Anthropic, and Fortune 500 enterprises like Siemens.</p><p style=\"min-height:1.5em\">Our team focuses on training small, efficient models that outperform much larger LLMs on domain-specific tasks (like OCR, structured extraction, tables). We move fast, prioritize practical results, and build tools that are open, reproducible, and built to last. You'll test hypotheses quickly, iterate on results, and balance experimental rigor with shipping to customers.</p><p style=\"min-height:1.5em\"></p><h2><strong>Day to day:</strong></h2><p style=\"min-height:1.5em\">A typical project might look like: identify a gap in extraction quality on long documents, train and benchmark a new model, optimize it for inference, and work with the team to ship it to users. Concretely:</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\"><strong>Train and evaluate models:</strong> Train task-specific models (OCR, layout, text recognition, extraction). Explore architectures and training strategies to optimize task performance. This includes our open source models, like Marker, Surya, and Chandra.</p></li><li><p style=\"min-height:1.5em\"><strong>Optimize inference:</strong> Profile and accelerate model inference across different hardware setups (H100s, B200s, L40s, CPUs).</p></li><li><p style=\"min-height:1.5em\"><strong>Ship to product:</strong> Work with the team to integrate models into our API and product, helping define how new capabilities surface for end users. You will be involved from model training through integration, although your work will be weighted much more towards the model side than the product side.</p></li><li><p style=\"min-height:1.5em\"><strong>Create and maintain datasets:</strong> Source, design, and clean datasets for supervised and synthetic training; create reproducible pipelines for data versioning and evaluation.</p></li><li><p style=\"min-height:1.5em\"><strong>Experiment and benchmark:</strong> Run ablations, track metrics, and publish findings that inform model design and internal research direction.</p></li><li><p style=\"min-height:1.5em\"><strong>Engage with users and partners:</strong> Occasionally join calls or Slack threads to better understand customer needs and inform your work.</p></li></ul><p style=\"min-height:1.5em\"></p><h2><strong>Ideal Candidate</strong></h2><p style=\"min-height:1.5em\">You've shipped models that made it into production. You understand how to balance exploration with delivery, and how to turn research insights into products people actually use.</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">3+ years experience training, fine-tuning, and evaluating deep learning models</p></li><li><p style=\"min-height:1.5em\">Trained at least one production-grade model or system used in real-world applications</p></li><li><p style=\"min-height:1.5em\">Deep expertise in PyTorch and Python, with strong fundamentals in deep learning (optimization, evaluation, architecture design)</p></li><li><p style=\"min-height:1.5em\">Comfortable with data engineering, benchmarking, and performance profiling across hardware setups</p></li><li><p style=\"min-height:1.5em\">Comfortable with an early stage startup - balance running ablations/benchmarks with shipping velocity</p></li></ul><p style=\"min-height:1.5em\"><strong>Bonus points if you:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Have experience with OCR, document AI, or structured extraction</p></li><li><p style=\"min-height:1.5em\">Have published work, whether that's a paper, a benchmark report, or a deep technical blog post</p></li><li><p style=\"min-height:1.5em\">Have been a major contributor to open-source projects, especially in ML, vision, or NLP</p></li><li><p style=\"min-height:1.5em\">Enjoy writing about your work and sharing learnings with the community</p></li></ul><p style=\"min-height:1.5em\"></p><h2><strong>Interview process</strong></h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">A 30-minute video call to evaluate fit</p></li><li><p style=\"min-height:1.5em\">90-minute live architecture discussion</p></li><li><p style=\"min-height:1.5em\">Culture fit interview/team meeting</p></li></ul><p style=\"min-height:1.5em\">At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.</p><p style=\"min-height:1.5em\"></p><div style=\"min-height:1.2em;margin-top:0;margin-bottom:0\"> </div><h1>Equal opportunity</h1><p style=\"min-height:1.5em\">Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.</p><p style=\"min-height:1.5em\">If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.</p>","descriptionPlain":"Salary range - $275k - $325k | Equity - 0.25% | In-person NYC\n\n\n\n\nABOUT DATALAB\n\nDatalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.\n\nWe’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.\n\n \n\n\nROLE OVERVIEW\n\nWe're looking for a Research Engineer to own problems end to end across our models, inference service, and product. You won't just train a model and hand it off. You'll take it from training through benchmarking, into our inference stack, and work with the team to integrate it into our products.\n\nWe're a small team that has shipped the current state of the art OCR model, Chandra. Our models collectively have 70k+ Github stars. Our tools are used internally at frontier AI labs like Anthropic, and Fortune 500 enterprises like Siemens.\n\nOur team focuses on training small, efficient models that outperform much larger LLMs on domain-specific tasks (like OCR, structured extraction, tables). We move fast, prioritize practical results, and build tools that are open, reproducible, and built to last. You'll test hypotheses quickly, iterate on results, and balance experimental rigor with shipping to customers.\n\n\n\n\nDAY TO DAY:\n\nA typical project might look like: identify a gap in extraction quality on long documents, train and benchmark a new model, optimize it for inference, and work with the team to ship it to users. Concretely:\n\n - Train and evaluate models: Train task-specific models (OCR, layout, text recognition, extraction). Explore architectures and training strategies to optimize task performance. This includes our open source models, like Marker, Surya, and Chandra.\n\n - Optimize inference: Profile and accelerate model inference across different hardware setups (H100s, B200s, L40s, CPUs).\n\n - Ship to product: Work with the team to integrate models into our API and product, helping define how new capabilities surface for end users. You will be involved from model training through integration, although your work will be weighted much more towards the model side than the product side.\n\n - Create and maintain datasets: Source, design, and clean datasets for supervised and synthetic training; create reproducible pipelines for data versioning and evaluation.\n\n - Experiment and benchmark: Run ablations, track metrics, and publish findings that inform model design and internal research direction.\n\n - Engage with users and partners: Occasionally join calls or Slack threads to better understand customer needs and inform your work.\n\n\n\n\nIDEAL CANDIDATE\n\nYou've shipped models that made it into production. You understand how to balance exploration with delivery, and how to turn research insights into products people actually use.\n\n - 3+ years experience training, fine-tuning, and evaluating deep learning models\n\n - Trained at least one production-grade model or system used in real-world applications\n\n - Deep expertise in PyTorch and Python, with strong fundamentals in deep learning (optimization, evaluation, architecture design)\n\n - Comfortable with data engineering, benchmarking, and performance profiling across hardware setups\n\n - Comfortable with an early stage startup - balance running ablations/benchmarks with shipping velocity\n\nBonus points if you:\n\n - Have experience with OCR, document AI, or structured extraction\n\n - Have published work, whether that's a paper, a benchmark report, or a deep technical blog post\n\n - Have been a major contributor to open-source projects, especially in ML, vision, or NLP\n\n - Enjoy writing about your work and sharing learnings with the community\n\n\n\n\nINTERVIEW PROCESS\n\n - A 30-minute video call to evaluate fit\n\n - 90-minute live architecture discussion\n\n - Culture fit interview/team meeting\n\nAt this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.\n\n\n\n \n\n\nEQUAL OPPORTUNITY\n\nDatalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.\n\nIf you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to."},{"id":"e9926a99-c5d2-40db-b6fd-4ca530e2a0b4","title":"Founding Customer Success Manager","department":"Sales","team":"Sales","employmentType":"FullTime","location":"New York Office","secondaryLocations":[],"publishedAt":"2026-07-23T21:49:56.294+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"New York","addressCountry":"United States","addressLocality":"New York"}},"jobUrl":"https://jobs.ashbyhq.com/datalab/e9926a99-c5d2-40db-b6fd-4ca530e2a0b4","applyUrl":"https://jobs.ashbyhq.com/datalab/e9926a99-c5d2-40db-b6fd-4ca530e2a0b4/application","descriptionHtml":"<p style=\"min-height:1.5em\"><strong>Base</strong> — $170k – $190k | <strong>OTE</strong> — ~$204k – $228k | <strong>Equity</strong> — 0.125% | <strong>In-person NYC</strong></p><p style=\"min-height:1.5em\"></p><h2>About Datalab</h2><p style=\"min-height:1.5em\">Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.</p><p style=\"min-height:1.5em\">We're at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, chandra, surya, marker, and lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.</p><p style=\"min-height:1.5em\"></p><h2>Role Overview</h2><p style=\"min-height:1.5em\">We're looking for a Founding Customer Success Manager to own the relationship with our enterprise customers after they sign. You'll be accountable for adoption, retention, and growth — making sure customers get real value from our models, stay for the long run, and expand their usage over time. You'll be the face of Datalab for every account you own: their trusted advisor, their first call when something matters, and the person keeping their goals moving forward.</p><p style=\"min-height:1.5em\">You will be the commercial and relationship manager who owns the account strategy after a deal has been signed. You'll orchestrate the right people internally so the customer always feels progress. You'll know the product and the customer's architecture well enough to lead most conversations yourself, and you'll pull in Engineering when an account needs it.</p><p style=\"min-height:1.5em\">We also have a long tail of self-serve customers using our API. Many are strong candidates to expand into enterprise contracts, and you'll own that self-serve → enterprise motion end to end — from spotting high-potential accounts to closing the upgraded contract.</p><p style=\"min-height:1.5em\">This role is ideal for someone who thrives at the intersection of customer strategy and technical fluency. You should be comfortable leading executive conversations, working with Engineering teams, be credible talking through an integration, and accountable for commercial outcomes.</p><p style=\"min-height:1.5em\"></p><p style=\"min-height:1.5em\"><strong>Day to day, you will:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Own the post-sales relationship for a portfolio of enterprise accounts — the trusted advisor and first point of escalation, orchestrating internal resources to keep customers delighted and unblocked.</p></li><li><p style=\"min-height:1.5em\">Lead onboarding, drive adoption, and execute technical delivery so customers reach first value fast and keep expanding usage from there.</p></li><li><p style=\"min-height:1.5em\">Run the cadence that builds trust: regular check-ins, business reviews, and proactive, continuous value delivery tied to each customer's goals.</p></li><li><p style=\"min-height:1.5em\">Own renewals and expansion, forecasting your book and turning healthy usage into larger commitments.</p></li><li><p style=\"min-height:1.5em\">Define and drive the self-serve → enterprise expansion strategy, spotting high-potential API accounts and converting them into contracts.</p></li><li><p style=\"min-height:1.5em\">Be the voice of the customer internally, feeding usage patterns and requests back to product and engineering to shape the roadmap.</p></li><li><p style=\"min-height:1.5em\">Build the repeatable playbooks (e.g., onboarding checklists, health scoring, QBR templates, expansion plays) that let account management scale.</p><p style=\"min-height:1.5em\"></p></li></ul><h2>How the role works with the team</h2><p style=\"min-height:1.5em\">You're the constant thread through a set of our major enterprise accounts. Our Solutions Engineer handles technical proof — evals, benchmarks, POCs, and getting customers live — and engineering owns the product and roadmap. You own the relationship and the outcome, pulling in the right people at the right moment. A typical thread looks like this:</p><p style=\"min-height:1.5em\">You’re reviewing account health and scanning for growth signals, and notice a customers API usage has spiked recently. You reach out and learn that they’re experimenting with a POC of a new tool that uses Datalab. You work with. the customer to understand their use case and timelines of this new POC. You loop in the Solutions Engineer, who helps build a benchmark tailored to their new use case. From this experience, you build an expansion playbook to identify and support API customers for this particular use case.</p><p style=\"min-height:1.5em\">The division: the SE runs the technical proof, engineering owns the product, and you own the relationship, the commercial outcome, and the coordination that ties it together.</p><p style=\"min-height:1.5em\">On the sales side: Our Account Executive owns new-logo deals and hands off at signature. You own the commercial relationship after signature/conversion: expansion, renewals, and converting self-serve API users into enterprise contracts, including the close. An existing self-serve customer who raises their hand to grow is your warm expansion opportunity.</p><p style=\"min-height:1.5em\"></p><h2>Ideal Candidate</h2><p style=\"min-height:1.5em\">You have experience owning customer relationships and a technical background that lets you go deep when it counts. You're high ownership, extremely organized, and excited about breaking down messy problems and turning them into consistent, repeatable processes.</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Experience in account management, customer success, or a technical customer-facing role, ideally owning renewals and expansion for enterprise accounts.</p></li><li><p style=\"min-height:1.5em\">Strong problem-solving skills and technical acumen — enough to navigate complex technical environments and map our capabilities to what a customer actually needs.</p></li><li><p style=\"min-height:1.5em\">Excellent written and verbal communication. You can translate technical concepts clearly for audiences ranging from engineers to executives.</p></li><li><p style=\"min-height:1.5em\">Comfort owning commercial outcomes: forecasting, negotiating renewals, and driving expansion.</p></li></ul><p style=\"min-height:1.5em\"><strong>Bonus points if you:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Have experience at an early-stage startup.</p></li><li><p style=\"min-height:1.5em\">Have a technical degree (CS, engineering, math, physics) even if you took a non-technical career path.</p></li><li><p style=\"min-height:1.5em\">Have familiarity with document AI, OCR, APIs, or developer-focused products.</p></li></ul><p style=\"min-height:1.5em\"></p><h2>Interview process</h2><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">A 30-minute video call to evaluate fit</p></li><li><p style=\"min-height:1.5em\">A 1-hour conversation to go deep on your past work</p></li><li><p style=\"min-height:1.5em\">A paid take-home project (~4-5 hours, $300) — design the post-sale motion for a segment of our customers: onboarding, health scoring, and an expansion play for the segment</p></li><li><p style=\"min-height:1.5em\">A 1-hour follow-up to discuss the project and meet the team</p></li></ul><p style=\"min-height:1.5em\">At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.</p><p style=\"min-height:1.5em\"></p><h1>Equal opportunity</h1><p style=\"min-height:1.5em\">Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.</p><p style=\"min-height:1.5em\">If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.</p>","descriptionPlain":"Base — $170k – $190k | OTE — ~$204k – $228k | Equity — 0.125% | In-person NYC\n\n\n\n\nABOUT DATALAB\n\nDatalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.\n\nWe're at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, chandra, surya, marker, and lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.\n\n\n\n\nROLE OVERVIEW\n\nWe're looking for a Founding Customer Success Manager to own the relationship with our enterprise customers after they sign. You'll be accountable for adoption, retention, and growth — making sure customers get real value from our models, stay for the long run, and expand their usage over time. You'll be the face of Datalab for every account you own: their trusted advisor, their first call when something matters, and the person keeping their goals moving forward.\n\nYou will be the commercial and relationship manager who owns the account strategy after a deal has been signed. You'll orchestrate the right people internally so the customer always feels progress. You'll know the product and the customer's architecture well enough to lead most conversations yourself, and you'll pull in Engineering when an account needs it.\n\nWe also have a long tail of self-serve customers using our API. Many are strong candidates to expand into enterprise contracts, and you'll own that self-serve → enterprise motion end to end — from spotting high-potential accounts to closing the upgraded contract.\n\nThis role is ideal for someone who thrives at the intersection of customer strategy and technical fluency. You should be comfortable leading executive conversations, working with Engineering teams, be credible talking through an integration, and accountable for commercial outcomes.\n\n\n\nDay to day, you will:\n\n - Own the post-sales relationship for a portfolio of enterprise accounts — the trusted advisor and first point of escalation, orchestrating internal resources to keep customers delighted and unblocked.\n\n - Lead onboarding, drive adoption, and execute technical delivery so customers reach first value fast and keep expanding usage from there.\n\n - Run the cadence that builds trust: regular check-ins, business reviews, and proactive, continuous value delivery tied to each customer's goals.\n\n - Own renewals and expansion, forecasting your book and turning healthy usage into larger commitments.\n\n - Define and drive the self-serve → enterprise expansion strategy, spotting high-potential API accounts and converting them into contracts.\n\n - Be the voice of the customer internally, feeding usage patterns and requests back to product and engineering to shape the roadmap.\n\n - Build the repeatable playbooks (e.g., onboarding checklists, health scoring, QBR templates, expansion plays) that let account management scale.\n   \n   \n\n\nHOW THE ROLE WORKS WITH THE TEAM\n\nYou're the constant thread through a set of our major enterprise accounts. Our Solutions Engineer handles technical proof — evals, benchmarks, POCs, and getting customers live — and engineering owns the product and roadmap. You own the relationship and the outcome, pulling in the right people at the right moment. A typical thread looks like this:\n\nYou’re reviewing account health and scanning for growth signals, and notice a customers API usage has spiked recently. You reach out and learn that they’re experimenting with a POC of a new tool that uses Datalab. You work with. the customer to understand their use case and timelines of this new POC. You loop in the Solutions Engineer, who helps build a benchmark tailored to their new use case. From this experience, you build an expansion playbook to identify and support API customers for this particular use case.\n\nThe division: the SE runs the technical proof, engineering owns the product, and you own the relationship, the commercial outcome, and the coordination that ties it together.\n\nOn the sales side: Our Account Executive owns new-logo deals and hands off at signature. You own the commercial relationship after signature/conversion: expansion, renewals, and converting self-serve API users into enterprise contracts, including the close. An existing self-serve customer who raises their hand to grow is your warm expansion opportunity.\n\n\n\n\nIDEAL CANDIDATE\n\nYou have experience owning customer relationships and a technical background that lets you go deep when it counts. You're high ownership, extremely organized, and excited about breaking down messy problems and turning them into consistent, repeatable processes.\n\n - Experience in account management, customer success, or a technical customer-facing role, ideally owning renewals and expansion for enterprise accounts.\n\n - Strong problem-solving skills and technical acumen — enough to navigate complex technical environments and map our capabilities to what a customer actually needs.\n\n - Excellent written and verbal communication. You can translate technical concepts clearly for audiences ranging from engineers to executives.\n\n - Comfort owning commercial outcomes: forecasting, negotiating renewals, and driving expansion.\n\nBonus points if you:\n\n - Have experience at an early-stage startup.\n\n - Have a technical degree (CS, engineering, math, physics) even if you took a non-technical career path.\n\n - Have familiarity with document AI, OCR, APIs, or developer-focused products.\n\n\n\n\nINTERVIEW PROCESS\n\n - A 30-minute video call to evaluate fit\n\n - A 1-hour conversation to go deep on your past work\n\n - A paid take-home project (~4-5 hours, $300) — design the post-sale motion for a segment of our customers: onboarding, health scoring, and an expansion play for the segment\n\n - A 1-hour follow-up to discuss the project and meet the team\n\nAt this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.\n\n\n\n\nEQUAL OPPORTUNITY\n\nDatalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.\n\nIf you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to."},{"id":"66231ee5-be4a-46ee-abd5-7b4bc65fbdea","title":"Founding Marketer","department":"Marketing","team":"Marketing","employmentType":"FullTime","location":"New York Office","secondaryLocations":[],"publishedAt":"2026-08-27T11:31:57.315+00:00","isListed":true,"isRemote":true,"workplaceType":"Hybrid","address":{"postalAddress":{"addressRegion":"New York","addressCountry":"United States","addressLocality":"New York"}},"jobUrl":"https://jobs.ashbyhq.com/datalab/66231ee5-be4a-46ee-abd5-7b4bc65fbdea","applyUrl":"https://jobs.ashbyhq.com/datalab/66231ee5-be4a-46ee-abd5-7b4bc65fbdea/application","descriptionHtml":"<h2><strong>About Datalab</strong></h2><p style=\"min-height:1.5em\">Datalab trains models that read documents reliably at scale. The world’s most important information is trapped in PDFs, scans, and files that can’t easily be parsed, and getting it out correctly matters. From frontier AI labs like Anthropic to Fortune 500s like Siemens, Datalab is where businesses turn when extraction has to be right.</p><p style=\"min-height:1.5em\">We hit an 8-figure run rate with a team of 7. We have hundreds of customers across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools - chandra, surya, marker, and lift - have 70,000+ GitHub stars, millions of monthly downloads, and broad developer mindshare. We’re backed by founding members of OpenAI, FAIR, and Hugging Face.</p><p style=\"min-height:1.5em\"></p><h2><strong>Role Overview</strong></h2><p style=\"min-height:1.5em\">Marketing at Datalab already has real infrastructure behind it. We run a combined launch and content calendar with a defined playbook - channels, owners, and cadence for every release - and we ship two to three launches a week against it. No one here works on marketing full-time, so nobody owns the larger question of how we show up when a developer, a buyer, or an AI agent goes looking for a document parser.</p><p style=\"min-height:1.5em\">You’d be our first dedicated marketing hire. You’d inherit a working system rather than a blank page, and the job is to make it run reliably, then make it considerably more ambitious. The natural place to build from is our open source work, which reaches a wide audience and currently does less for us than it could.</p><p style=\"min-height:1.5em\">This is a hands-on role. On a team of 7, you’ll write the blog post, draft the tweets, restructure the docs page, brief the agency if we hire one, and pull the numbers yourself. We’re looking for someone who has done this work somewhere it was done well, and who wants to keep doing it rather than direct someone else who does.</p><p style=\"min-height:1.5em\">There is strong potential for this to grow into a leadership role as we hire more people into the marketing team.</p><p style=\"min-height:1.5em\">An increasing share of our buyers never touch a search results page; they ask a model. We want someone who takes seriously how Datalab gets recommended by LLMs and agents - how our site, docs, and benchmarks are structured, where models pull their answers from, and how any of it gets measured. Nobody has fully figured this out yet, and we’d like you to be the person who works it out here.</p><p style=\"min-height:1.5em\">If you want to inherit a brand book, a content team, and a defined category, this isn’t the right role. If you want to shape how a technical category gets talked about, with real customers and real distribution to work from, it’s a good one.</p><p style=\"min-height:1.5em\"></p><h2><strong>Day to day you will:</strong></h2><p style=\"min-height:1.5em\">A typical week might look like: shipping the launch post for a new open source model, ghostwriting a thread for one of our researchers, editing a vertical explainer on tracked-changes extraction for legal teams, restructuring a docs page so a model can actually cite it, answering a sharp question in Discord, and digging into why organic signups dipped last week.</p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\"><strong>Own our model launches, especially open source.</strong> Every release - chandra, surya, marker, lift, and what comes next - should land. You’ll build the assets (launch posts, benchmark write-ups, demos, docs), coordinate the team’s posts, work HN and Reddit and the ML community, handle press where it’s worth handling, and run the retro afterward so the next one lands harder.</p></li><li><p style=\"min-height:1.5em\"><strong>Take over and scale the content engine.</strong> Inherit the calendar, tighten the cadence, and own the strategy and much of the writing. Educational content for developers evaluating us, and vertical-specific content for the buyers who need to see their own documents in the example - legal, healthcare, financial services, insurance, logistics, government.</p></li><li><p style=\"min-height:1.5em\"><strong>Ghostwrite for the team.</strong> Our engineers and researchers know things worth writing about and mostly won’t write it themselves. You’ll draft in their voice, in a way they’re glad to put their name on, and turn the team into a distribution channel.</p></li><li><p style=\"min-height:1.5em\"><strong>Make Datalab legible to LLMs and agents.</strong> Structure the site, docs, and benchmark content so that models recommend us when someone asks how to parse a PDF. Own the technical work behind it, figure out what “ranking” even means in this context, and build a way to measure whether it’s working.</p></li><li><p style=\"min-height:1.5em\"><strong>Treat the website as our front door.</strong> Most of our revenue starts with someone landing on <a target=\"_blank\" rel=\"noopener noreferrer nofollow\" href=\"http://datalab.to\">datalab.to</a>. Own the narrative, structure, and conversion path of the site and the top of our docs, working with engineering to ship changes.</p></li><li><p style=\"min-height:1.5em\"><strong>Own the demand channels and the budget behind them.</strong> SEO, paid, landing pages, conversion. Propose and manage the spend, and own the numbers end to end — partnering with our Business Operations hire on funnel reporting and attribution so we know which channels actually produce revenue.</p></li><li><p style=\"min-height:1.5em\"><strong>Arm the sales team.</strong> Build the case studies, benchmark comparisons, one-pagers, and decks our AE and solutions engineers need in live deals. Our strongest proof points are customers with hard extraction problems; turn those into assets that help close business and double as vertical content.</p></li><li><p style=\"min-height:1.5em\"><strong>Show up in the community.</strong> Our Discord and GitHub are the largest audience we own. Set the tone there, spot what people keep asking for, and turn recurring questions into documentation and content. Represent us at conferences, meetups, and on podcasts — and build the speaking and events motion, including getting our researchers in front of the right rooms.</p></li><li><p style=\"min-height:1.5em\"><strong>Track where AI is going and translate it into positioning.</strong> Agents, evals, context engineering, the shifting shape of the document AI market. Bring us a point of view on what it means for how we talk about ourselves, not just a summary of the news.</p></li><li><p style=\"min-height:1.5em\"><strong>Feed the loop back to product and sales.</strong> What messaging converts, what objections keep surfacing, what the community keeps asking for — bring structured signal to the founder and the GTM team on a regular cadence.</p></li></ul><h2><strong>What success looks like</strong></h2><p style=\"min-height:1.5em\"><strong>By month 3</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">You’ve owned at least one model launch end to end, and it outperformed the last comparable release on a metric we agreed on in advance.</p></li><li><p style=\"min-height:1.5em\">The launch and content calendar is shipping on schedule without slipping, with the team contributing under their own names.</p></li><li><p style=\"min-height:1.5em\">Baselines are in place: organic traffic, signups by source, LLM-recommendation visibility, and a first read on what’s working.</p></li></ul><p style=\"min-height:1.5em\"><strong>By month 12</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Inbound is meaningfully up and attributable, with at least one channel you built from nothing producing consistent pipeline.</p></li><li><p style=\"min-height:1.5em\">The launch playbook is documented well enough that a new hire could run a release with it.</p></li><li><p style=\"min-height:1.5em\">Datalab shows up when a model is asked how to parse documents at scale — and you can prove it moved because of work you did.</p></li></ul><h2><strong>Ideal candidate</strong></h2><p style=\"min-height:1.5em\">You’ve built this before at a company where it mattered. You write well and fast, and you’d rather ship a good post today than a perfect one next month. You’re comfortable being measured - you’d rather have a number attached to your work than not. You’re ambitious about where this goes and want more scope over time, but you’re not waiting for it before you start doing the work.</p><p style=\"min-height:1.5em\">You’re technical enough to hold your own. Our audience is engineers, ML leads, and CTOs, and they can tell instantly when marketing doesn’t understand the product. You should be able to read a benchmark table and know what it means, run our models yourself, and write something a skeptical developer on Hacker News would find useful rather than annoying.</p><p style=\"min-height:1.5em\">You’re collaborative and low-ego. You’ll be pulling engineers into launches, asking researchers for time, and putting your writing out under someone else’s name regularly. You make the call that’s right for the company, and you’re glad when the work lands even if the byline isn’t yours.</p><p style=\"min-height:1.5em\"><strong>We’re eager to work with someone who has:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">7+ years in marketing, with meaningful time owning content, launches, and demand generation for a technical or developer-facing product.</p></li><li><p style=\"min-height:1.5em\">Run launches at a high level - coordinated messaging, assets, press, and community across a real release cycle - and can point to both the results and the parts you personally shipped.</p></li><li><p style=\"min-height:1.5em\">Built a content program from strategy through published work, and written most of it yourself.</p></li><li><p style=\"min-height:1.5em\">Owned SEO and paid channels directly, including budget, with numbers on what moved and what didn’t.</p></li><li><p style=\"min-height:1.5em\">Operated in early-stage environments where the playbook didn’t exist and the team was small.</p></li><li><p style=\"min-height:1.5em\">Can use AI tools like Claude and Cowork effectively to streamline your work and build content systems.</p></li></ul><p style=\"min-height:1.5em\"><strong>Bonus points if you:</strong></p><ul style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">Have marketed an open source project or built inside a developer community (GitHub, HN, Discord, ML Twitter).</p></li><li><p style=\"min-height:1.5em\">Have worked on document AI, OCR, IDP, or data extraction - or adjacent infrastructure and developer tooling.</p></li><li><p style=\"min-height:1.5em\">Have run experiments in LLM/AI-assistant discoverability and have opinions about what actually works.</p></li><li><p style=\"min-height:1.5em\">Have ghostwritten for founders or technical leaders and have samples you can talk through.</p></li><li><p style=\"min-height:1.5em\">Are comfortable enough with design to get a launch asset out the door yourself, and know when to bring in help.</p></li><li><p style=\"min-height:1.5em\">Have marketed into regulated or public-sector buyers, where the proof burden is higher.</p></li><li><p style=\"min-height:1.5em\">Have hired contractors, agencies, or freelancers and gotten good work out of them.</p></li><li><p style=\"min-height:1.5em\">Want to build and lead a marketing team, and can tell us why you think you’re ready for that next.</p></li><li><p style=\"min-height:1.5em\">Are active in the AI community and actually use open source models and tools.</p></li></ul><h2><strong>Interview process</strong></h2><ol style=\"min-height:1.5em\"><li><p style=\"min-height:1.5em\">A 30-minute video call with the founder to evaluate fit.</p></li><li><p style=\"min-height:1.5em\">A 1-hour conversation going deep on a launch or content program you’ve owned end to end - what you did, what worked, what you’d change.</p></li><li><p style=\"min-height:1.5em\">A 2-hour onsite in NYC, working through three things: traffic and funnel data, a live editing and ghostwriting session using our actual drafts and voices, and a discussion of a short strategy brief.</p></li><li><p style=\"min-height:1.5em\">Final conversations with the team.</p></li></ol><p style=\"min-height:1.5em\">We expect you to use AI tools throughout, including on the prepared brief - we use them constantly and want to see how you work, not watch you pretend you don’t.</p><p style=\"min-height:1.5em\">At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.</p><p style=\"min-height:1.5em\"></p><h1>Equal opportunity</h1><p style=\"min-height:1.5em\">Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.</p><p style=\"min-height:1.5em\">If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.</p>","descriptionPlain":"ABOUT DATALAB\n\nDatalab trains models that read documents reliably at scale. The world’s most important information is trapped in PDFs, scans, and files that can’t easily be parsed, and getting it out correctly matters. From frontier AI labs like Anthropic to Fortune 500s like Siemens, Datalab is where businesses turn when extraction has to be right.\n\nWe hit an 8-figure run rate with a team of 7. We have hundreds of customers across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools - chandra, surya, marker, and lift - have 70,000+ GitHub stars, millions of monthly downloads, and broad developer mindshare. We’re backed by founding members of OpenAI, FAIR, and Hugging Face.\n\n\n\n\nROLE OVERVIEW\n\nMarketing at Datalab already has real infrastructure behind it. We run a combined launch and content calendar with a defined playbook - channels, owners, and cadence for every release - and we ship two to three launches a week against it. No one here works on marketing full-time, so nobody owns the larger question of how we show up when a developer, a buyer, or an AI agent goes looking for a document parser.\n\nYou’d be our first dedicated marketing hire. You’d inherit a working system rather than a blank page, and the job is to make it run reliably, then make it considerably more ambitious. The natural place to build from is our open source work, which reaches a wide audience and currently does less for us than it could.\n\nThis is a hands-on role. On a team of 7, you’ll write the blog post, draft the tweets, restructure the docs page, brief the agency if we hire one, and pull the numbers yourself. We’re looking for someone who has done this work somewhere it was done well, and who wants to keep doing it rather than direct someone else who does.\n\nThere is strong potential for this to grow into a leadership role as we hire more people into the marketing team.\n\nAn increasing share of our buyers never touch a search results page; they ask a model. We want someone who takes seriously how Datalab gets recommended by LLMs and agents - how our site, docs, and benchmarks are structured, where models pull their answers from, and how any of it gets measured. Nobody has fully figured this out yet, and we’d like you to be the person who works it out here.\n\nIf you want to inherit a brand book, a content team, and a defined category, this isn’t the right role. If you want to shape how a technical category gets talked about, with real customers and real distribution to work from, it’s a good one.\n\n\n\n\nDAY TO DAY YOU WILL:\n\nA typical week might look like: shipping the launch post for a new open source model, ghostwriting a thread for one of our researchers, editing a vertical explainer on tracked-changes extraction for legal teams, restructuring a docs page so a model can actually cite it, answering a sharp question in Discord, and digging into why organic signups dipped last week.\n\n - Own our model launches, especially open source. Every release - chandra, surya, marker, lift, and what comes next - should land. You’ll build the assets (launch posts, benchmark write-ups, demos, docs), coordinate the team’s posts, work HN and Reddit and the ML community, handle press where it’s worth handling, and run the retro afterward so the next one lands harder.\n\n - Take over and scale the content engine. Inherit the calendar, tighten the cadence, and own the strategy and much of the writing. Educational content for developers evaluating us, and vertical-specific content for the buyers who need to see their own documents in the example - legal, healthcare, financial services, insurance, logistics, government.\n\n - Ghostwrite for the team. Our engineers and researchers know things worth writing about and mostly won’t write it themselves. You’ll draft in their voice, in a way they’re glad to put their name on, and turn the team into a distribution channel.\n\n - Make Datalab legible to LLMs and agents. Structure the site, docs, and benchmark content so that models recommend us when someone asks how to parse a PDF. Own the technical work behind it, figure out what “ranking” even means in this context, and build a way to measure whether it’s working.\n\n - Treat the website as our front door. Most of our revenue starts with someone landing on datalab.to http://datalab.to. Own the narrative, structure, and conversion path of the site and the top of our docs, working with engineering to ship changes.\n\n - Own the demand channels and the budget behind them. SEO, paid, landing pages, conversion. Propose and manage the spend, and own the numbers end to end — partnering with our Business Operations hire on funnel reporting and attribution so we know which channels actually produce revenue.\n\n - Arm the sales team. Build the case studies, benchmark comparisons, one-pagers, and decks our AE and solutions engineers need in live deals. Our strongest proof points are customers with hard extraction problems; turn those into assets that help close business and double as vertical content.\n\n - Show up in the community. Our Discord and GitHub are the largest audience we own. Set the tone there, spot what people keep asking for, and turn recurring questions into documentation and content. Represent us at conferences, meetups, and on podcasts — and build the speaking and events motion, including getting our researchers in front of the right rooms.\n\n - Track where AI is going and translate it into positioning. Agents, evals, context engineering, the shifting shape of the document AI market. Bring us a point of view on what it means for how we talk about ourselves, not just a summary of the news.\n\n - Feed the loop back to product and sales. What messaging converts, what objections keep surfacing, what the community keeps asking for — bring structured signal to the founder and the GTM team on a regular cadence.\n\n\nWHAT SUCCESS LOOKS LIKE\n\nBy month 3\n\n - You’ve owned at least one model launch end to end, and it outperformed the last comparable release on a metric we agreed on in advance.\n\n - The launch and content calendar is shipping on schedule without slipping, with the team contributing under their own names.\n\n - Baselines are in place: organic traffic, signups by source, LLM-recommendation visibility, and a first read on what’s working.\n\nBy month 12\n\n - Inbound is meaningfully up and attributable, with at least one channel you built from nothing producing consistent pipeline.\n\n - The launch playbook is documented well enough that a new hire could run a release with it.\n\n - Datalab shows up when a model is asked how to parse documents at scale — and you can prove it moved because of work you did.\n\n\nIDEAL CANDIDATE\n\nYou’ve built this before at a company where it mattered. You write well and fast, and you’d rather ship a good post today than a perfect one next month. You’re comfortable being measured - you’d rather have a number attached to your work than not. You’re ambitious about where this goes and want more scope over time, but you’re not waiting for it before you start doing the work.\n\nYou’re technical enough to hold your own. Our audience is engineers, ML leads, and CTOs, and they can tell instantly when marketing doesn’t understand the product. You should be able to read a benchmark table and know what it means, run our models yourself, and write something a skeptical developer on Hacker News would find useful rather than annoying.\n\nYou’re collaborative and low-ego. You’ll be pulling engineers into launches, asking researchers for time, and putting your writing out under someone else’s name regularly. You make the call that’s right for the company, and you’re glad when the work lands even if the byline isn’t yours.\n\nWe’re eager to work with someone who has:\n\n - 7+ years in marketing, with meaningful time owning content, launches, and demand generation for a technical or developer-facing product.\n\n - Run launches at a high level - coordinated messaging, assets, press, and community across a real release cycle - and can point to both the results and the parts you personally shipped.\n\n - Built a content program from strategy through published work, and written most of it yourself.\n\n - Owned SEO and paid channels directly, including budget, with numbers on what moved and what didn’t.\n\n - Operated in early-stage environments where the playbook didn’t exist and the team was small.\n\n - Can use AI tools like Claude and Cowork effectively to streamline your work and build content systems.\n\nBonus points if you:\n\n - Have marketed an open source project or built inside a developer community (GitHub, HN, Discord, ML Twitter).\n\n - Have worked on document AI, OCR, IDP, or data extraction - or adjacent infrastructure and developer tooling.\n\n - Have run experiments in LLM/AI-assistant discoverability and have opinions about what actually works.\n\n - Have ghostwritten for founders or technical leaders and have samples you can talk through.\n\n - Are comfortable enough with design to get a launch asset out the door yourself, and know when to bring in help.\n\n - Have marketed into regulated or public-sector buyers, where the proof burden is higher.\n\n - Have hired contractors, agencies, or freelancers and gotten good work out of them.\n\n - Want to build and lead a marketing team, and can tell us why you think you’re ready for that next.\n\n - Are active in the AI community and actually use open source models and tools.\n\n\nINTERVIEW PROCESS\n\n 1. A 30-minute video call with the founder to evaluate fit.\n\n 2. A 1-hour conversation going deep on a launch or content program you’ve owned end to end - what you did, what worked, what you’d change.\n\n 3. A 2-hour onsite in NYC, working through three things: traffic and funnel data, a live editing and ghostwriting session using our actual drafts and voices, and a discussion of a short strategy brief.\n\n 4. Final conversations with the team.\n\nWe expect you to use AI tools throughout, including on the prepared brief - we use them constantly and want to see how you work, not watch you pretend you don’t.\n\nAt this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.\n\n\n\n\nEQUAL OPPORTUNITY\n\nDatalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.\n\nIf you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to."}],"apiVersion":"1"}