Artificial Intelligence Engineer
🏢 UpDraft
🗓️ Publicada em: 08 de setembro de 2026, 13:25
Who We Are
Behind every 10-K, audit, and acquisition is accounting work that has to be timely, accurate, and auditable. Today that work means partners and senior accountants reading thousands of pages of contracts and guidance by hand, under deadline, with their name on the conclusion.
UpDraft is the AI work platform for corporate accounting, from cash rec to 10-K. We started with technical accounting: our agent reads the guidance, finds the right parts of the contract, applies the right rules, and drafts conclusions the way a technical accountant and an auditor would expect. Top-25 CPA firms, elite advisory practices, Fortune 1000 finance teams, and pre-IPO companies already trust it, and we're expanding the same recipe to the rest of accounting.
We're a talent-dense small team, VC-backed. Engineering is led by a technical co-founder, and everyone ships to production.
Why This Job Is Exciting
Are you an AI engineer who wants real impact on what a product becomes? We're flat and small: your opinion carries real weight and can genuinely drive what we do. This is a role on a product where the agent IS the product. The quality of our agentic systems is directly what customers pay for: how well they read a 500-page contract, pick and apply the right guidance, and write the result into the formal deliverable. Concretely, you will:
- Own agentic systems end to end. Design, build, and harden multi-step, tool-using agent loops that turn messy source documents into auditor-ready work product. You own reliability, observability, and failure modes, not just the happy path.
- Use evaluations pragmatically. Crafting agentic products means knowing when a change actually helped, which is hard when both the agents and the product keep shifting. You'll bring judgment about where evaluations earn their keep: curated datasets, offline replay, scorers and judges, and regression alerts where they pay off; targeted smoke tests and metrics where they don't; and no noise dressed up as rigor. We move fast with confidence.
- Design feedback loops from real usage. Collect, clean, and interpret user signals to inform model and harness changes, and build the analysis tooling to debug agent behavior: deep dives on failure modes, clustering themes, and surfacing actionable insights.
- Make quality measurable and operational. Improve reliability and guardrails by defining what good, bad, and degraded sessions look like, with the alerting and triage primitives to act on them.
- Push retrieval and context engineering. Contracts come with amendments, exhibits, and cross-references. You'll work on how agents ground themselves in the right documents and the right paragraphs of accounting guidance, drawing on our library of 500+ accounting guides and skills.
- Shape the product, not just the code. On a team this small, there is no layer between you, the founders, and the customer. What you learn from users goes straight into what we build next.
Month 1 is for ramping up. You'll learn the codebase, the product, our customers, and their use cases, and you'll ship while you learn: small and medium improvements, closed feature gaps, real code in production from week two.
Month 2, you'll lead your first initiatives end to end, sometimes solo, sometimes alongside teammates, from scoping through shipping.
Month 3, you'll have a real voice in the roadmap: bringing your own ideas, proposing new technologies worth trying, and shaping where the product goes next.
About You
Who you are
You're wired for early-stage startups, not big enterprises. You're the kind of person who goes deep on things: your work, a hobby, a rabbit hole you couldn't leave alone. Ambiguity and a bit of chaos energize you more than they stress you. Your career shows it: fast progression, ownership of problems nobody else wanted, and a track record other engineers respect and follow.
How you build
You've barely written code by hand in months. You run multiple coding agents in parallel, and your time goes to reviewing their output, nudging them in the right direction, and designing architectures they can execute against. You're a super user of these tools, not a casual one. And when a library doesn't do what you need and nothing else does either, you fork it, push a PR upstream, run your own version, or get the maintainers on a call and make the case for your feature on their roadmap.
What you know
- You've built LLM-powered systems for real, not a chatbot demo or an email assistant. You've dealt with context management, tool clearing, and compaction, and you know the small gotchas of the Anthropic and OpenAI APIs from experience.
- You've worked with retrieval and context engineering over long, structured documents.
- You've built evaluation and analytics for agentic systems before: eval datasets, scorers and judges, regression checks, plus the usage analysis side of digging into real sessions, finding failure modes, and turning them into fixes.
- You're fluent in Python and JavaScript. You don't need to know every corner of either language, but you know the best practices cold, and you're strong in Clean Code, Clean Architecture, and software architecture in general.
- You don't need an accounting background. You need curiosity about a deep domain and a bar for correctness that matches work that gets defended to auditors.
Nice to haves:
- Background in fintech, legal tech, or another correctness-critical domain
- Previous founder or founding-team experience
- Experience reading, editing, and rendering .docx and .xlsx files programmatically: our product works with these documents every day
Level and Compensation
This is a senior role. We hire remotely across LATAM, and we pay above local market. This is a high-demand craft and we want you focused on the work, not on your bills.
You'll also get equity, in stock options, with the exact numbers shared before any offer. We're early, which means real ownership and a real chance to grow with the company in scope, responsibility, and upside.
Four weeks of PTO per year.
And to be transparent about the trade: this is a fast-paced startup, and this job will be a big part of your life. In exchange, you'll work with a frontier AI stack, ship to demanding customers, and move at a pace that will change the trajectory of your career.
Interview Process
We designed the process to look like the actual job: no whiteboard, no puzzles, no writing algorithms from memory. Expect about 3.5 hours of live conversation plus one take-home, and we move fast between stages.
- Introduction stage - we get to know each other:
- [30m] Intro call. The company, the role, your expectations and ours.
- [60m] Career deep dive. We walk through your story: what you've owned, how you've grown, what you go deep on, and how you work with coding agents today.
- Technical stage - you show us how you actually work:
- [Async] Take-home exercise, about a day of work. A scoped, realistic task. We expect you to use coding agents the same way you would on the job, and to send a short write-up of your decisions and trade-offs.
- [90m] Walkthrough and live extension. You walk us through your take-home and defend your choices, then extend it live with your own tools and your normal workflow.
- Final stage - we close the loop:
- [30m] Co-founder conversation. Product, customers, and values with our co-founder.
- We check references and conduct your background check.
How to Apply
We don't review applications sent through LinkedIn. Apply through this form
It takes about ten minutes. We ask for your name, region, LinkedIn profile, and CV, plus one short question about something you've actually built or debugged (use whatever tools you want to draft your answer, but it needs to be yours. You'll defend every word of it live in the interview). We read every submission that comes through the form.