job description

AI Engineer

job description

AI Engineer

About Uplevyl 

Uplevyl builds AI-powered knowledge and community infrastructure for organizations serving women. Our products include UpGenie (our domain-specific AI assistant), WeHub (our community platform), and UpSocial (a social platform for women). We work with mission-driven partners to turn complex, high-stakes information into clear, trustworthy guidance, and to build systems that create lasting impact. 

We hold a simple conviction: in high-stakes domains like rights, law, and financial security, a generic AI is not enough. The answers people stake their lives and livelihoods on need a purpose-built system with verified, native data. That is what we build, and it is why the work is urgent. AI is reshaping how the world learns, works, and earns, and the women we serve cannot afford to be left off that train. 

We have also made a deliberate choice about how we build: a small team of exceptional people, paid well above market, each doing work that would normally take several. We would rather be ten people who move the world than thirty who move paper. That choice sets the bar for every hire, including this one.

About The Role

AI Engineer is the builder of UpGenie's core. You own the AI layer of a production system that answers high-stakes questions about rights, law, and safety, where the person on the other end may be staking a job or their wellbeing on what we return.

Our system retrieves across tens of thousands of jurisdictions and multiple legal domains, reasons over structured metadata like eligibility thresholds and enforcement agencies, validates its own sources, and proves its accuracy through evals. You will work on retrieval quality, schema design, evaluation infrastructure, and the roadmap from today's pipeline to a genuinely agentic system.

You will ship to production from your first weeks, with real users and real partners depending on what you build.

What You'll Own

  • Retrieval quality. The RAG architecture end to end: chunking, metadata filtering, ranking and re-ranking, and the retrieval logic that finds the right law among tens of thousands of jurisdictions across overlapping federal, state, county, and city layers.

  • The schema and metadata layer. Design the structured data that makes legal knowledge machine-usable: thresholds, agencies, citations, jurisdiction hierarchies. This layer is our moat, and you will shape it.

  • Evals as infrastructure. Build the evaluation framework that proves answer quality: hallucination traps, citation validation, cascade-logic tests across jurisdiction hierarchies, and statistically defensible coverage claims. If we say the system is right, you built the proof.

  • Source validation. Own the AI-first pipeline that decides which legal sources enter the system and which are rejected, and make its failure modes visible before they reach a user.

  • The agentic roadmap. Move us from hard-coded workflow logic to schema-aware prompting and tool-using agents. You will help define what V3 of the system looks like and then build it.

Who We're Looking For

We hire across levels. Working systems you have shipped matter more than years. We will ask you to build in our process, because we have learned that talking about AI and shipping AI are different skills.

  • You ship working code. Production systems you built, not notebooks you ran. You are fluent in Python and comfortable owning a service that real users hit.

  • Strong AI fundamentals. You understand retrieval, embeddings, context management, prompting, and evaluation at the level of mechanisms, and you know where these systems break. You use AI daily to build and you have formed your own view on what is real versus hype.

  • Rigor. In our domain, a confident wrong answer can harm someone. You test before you claim, you measure before you optimize, and you flag what you are unsure about.

  • High agency. You unblock yourself. Ambiguous problem, no owner in sight: you scope it, propose an approach, and build.

  • Range, picked up fast. You move across retrieval, data pipelines, prompting, and evaluation, and you learn a new tool in hours, not weeks.

  • A clear communicator. You can explain a retrieval failure to a non-technical partner and defend a design decision to a senior engineer, in plain words either way.

Even Better If

  • You have shipped RAG or LLM systems that are live in production today.

  • You have worked with Vertex AI, GCP, or comparable cloud AI stacks.

  • You have built evaluation or testing infrastructure for AI systems.

  • You have worked with legal, regulated, or otherwise high-stakes data.

  • You have public work we can look at: open source, writing, side projects that did real work for real users.

Before You Apply

We want to be direct about what this is. This is a high-bar role on a small, senior team. The work is demanding, the expectations are real, and the rewards, in compensation, in ownership, and in what you will learn, match them. If you want a defined lane and a slow ramp, this is not the right fit, and that is fine. If you want to build the AI core of a system where accuracy genuinely matters, and grow into the technical backbone of a small, ambitious team, we want to hear from you. Our interview process is thorough, and we will walk you through every step of it in our first conversation.

How We Work at Uplevyl

  • We finish what we start, and we do it well. We follow through on what we say we'll do, and we care about the difference our work makes. 

  • We listen before we decide. Before we act, we ask who it actually affects: a customer, a partner, a teammate, or the communities we serve. Trust here is built the plain way, by consistently showing up for people. 

  • We move with ownership and urgency. We don't wait for perfect information or for someone else to raise their hand. If something looks like it could go wrong, we say so early. We make the call, we move fast, and we hold a high bar for quality. 

  • We're better together than alone. We work across teams, say what we actually think, and go out of our way to help each other succeed. Leadership here is measured by the impact you create. 

  • We stay curious and keep learning, including how to use AI well. We hold ourselves to the bar we're building toward: questioning our own assumptions and using AI to think better and move faster.