AI Model Training Quality Engineer
1 day ago
Portland
Job DescriptionAbout Sapience AI Sapience AI is the collective intelligence platform for professional communities. We sit above the CRMs, AMS platforms, and knowledge bases that organizations already run, and we turn the expertise scattered across them into something every member can search, act on, and share. The intelligence a community needs is already inside it. Most organizations just cannot reach it. Knowledge lives in silos, in legacy systems, in the heads of a few experts, and in fragmented records no one can connect. We change that. Our work is grounded in four commitments: technology elevates people and never replaces them, the best expertise is already inside the community, everything is built on trust, and every deployment is purpose-driven for the organization it serves. Let's achieve more, together. Where this role sits This role owns the quality of the models at the core of Sapience AI. You define, lead, and perform the training, reinforcement, and evaluation that turn open-weight foundation models into systems that reason well over a community's knowledge and earn members' trust. You work where research, data, and engineering meet: training and fine-tuning open model weights, reinforcing them toward the behavior communities need, and holding a rigorous evaluation bar that says, with evidence, whether a model is good enough to ship. Your work feeds directly into the COGENT architecture and the MINERVA platform. You are both a leader and a practitioner. You set the standard for what model quality means at Sapience AI, and you do the hands-on training, reinforcement, and evaluation that meet it. Why this role exists Collective intelligence is only as trustworthy as the models beneath it. If a model reasons poorly, drifts from the truth, or cannot be evaluated honestly, every system built on it inherits that weakness, and the trust communities place in the platform erodes. Training and reinforcing open-weight models well is a discipline in its own right: choosing the right approach, curating the right data, running reinforcement that moves behavior in the right direction, and measuring quality rigorously enough to know it worked. Done poorly, it produces confident, unmeasured guesses. Done well, it produces models people can rely on. The AI Model Training Quality Engineer owns that discipline end to end. You define what good looks like, lead the training and reinforcement that get there, and hold the evaluation bar that keeps Sapience AI's models honest, capable, and safe. What you will own (Areas of Responsibility) You hold eight areas of responsibility across model training and quality. Each one is yours to define, lead, and perform. 1. Model training strategy and quality standards • Define what model quality means at Sapience AI: the capabilities, groundedness, calibration, and safety a model must meet before it ships., • Own the training strategy for open-weight foundation models, including when to continue pretraining, fine-tune, distill, or adapt., • Set the standards, gates, and reproducibility practices that make training decisions defensible and repeatable.2. Open-weight model training and fine-tuning, • Train and fine-tune open-weight models, including supervised fine-tuning, domain adaptation, and instruction tuning for the needs of professional communities., • Run efficient training at scale, using parameter-efficient methods where they fit and full fine-tuning where they do not., • Make sound trade-offs across capability, cost, latency, and the constraints of production.3. Reinforcement and preference optimization, • Lead reinforcement and alignment work, including RLHF, RLAIF, and direct preference methods, to move model behavior toward what communities actually need., • Design and manage preference and reward data, and the reward or preference signals that shape behavior., • Reinforce toward groundedness, honesty about uncertainty, and safety, not just fluency or benchmark gains.4. Evaluation, benchmarking, and quality measurement, • Build and own the evaluation that decides whether a model is good enough: accuracy, groundedness, calibration, safety, and robustness., • Design evaluations that reflect real community needs, not just public benchmarks, and cover both offline tests and online behavior., • Run rigorous comparisons across models and training runs, and report results honestly, including where a model falls short.5. Training and evaluation data, • Partner with data and knowledge engineering on the training, preference, and evaluation datasets that quality depends on., • Own data quality for training and evaluation, including contamination control, deduplication, coverage, and bias., • Ground training and evaluation in the KO graph where it strengthens reasoning over community knowledge.6. Trust, safety, and responsible model behavior, • Reduce the failure modes that erode trust, including hallucination, confident errors, and bias, and improve calibration., • Build model behavior that handles uncertainty honestly and stays within safe bounds., • Treat safety and trust as part of training and evaluation, not a later review.7. Productionization and regression prevention, • Take trained and reinforced models to production in partnership with ML infrastructure and applied AI., • Own the quality gates and regression tests that stop a worse model from shipping., • Monitor model quality in production and close the loop when behavior drifts.8. Leadership, standards, and team enablement, • Set the training-quality bar for the organization and lead others to meet it., • Make training and evaluation reproducible, documented, and reusable so results can be trusted and built on., • Mentor engineers and researchers, and raise the standard of how the whole team trains and evaluates models.AI-augmented ways of working AI is both your subject and your tool. You use AI to accelerate data curation, generate and grade candidate outputs, scale evaluation, and reason about results, while you own the training decisions, the quality bar, and the safety judgments that AI cannot make for you. The standard is human in partnership: AI accelerates the work, you own the judgment, the interpretation, and the call. The people who create the most value here are not the ones producing the most output. They are the ones turning evidence into models people can trust. What this role is not To keep the boundary clear: • This is not a pure research role. You are accountable for models that ship and hold up in production, not only for publications or prototypes., • This is not an inference-infrastructure role. You partner with ML infrastructure on serving, but your focus is training, reinforcement, and quality., • This is not a data-engineering-only role. You partner with data teams on datasets; you own the quality of training, reinforcement, and evaluation., • This is not an agent or applied-AI role. You own the model itself; applied AI builds the agents and behavior on top of it., • This is not a benchmark-chasing role. You are measured on trustworthy, capable models in real community settings, not leaderboard scores alone.What success looks like We measure this role on the quality and trustworthiness of the models it produces: • Better models. Trained and reinforced models reason more capably and more reliably over community knowledge., • Honest evaluation. The organization has a rigorous, trustworthy picture of what each model can and cannot do., • Grounded and calibrated. Confident errors go down, calibration improves, and answers are more traceable., • Safe behavior. Models handle uncertainty honestly and stay within safe bounds., • Reproducible training. Training and evaluation are documented and repeatable, so results can be trusted and built on., • No silent regressions. Quality gates stop worse models from shipping, and drift is caught in production., • A stronger bar. The whole team trains and evaluates to a higher, clearer standard because of your leadership.Who you areRequired qualifications, • Five or more years in machine learning, with strong hands-on experience training and fine-tuning modern models., • Direct experience training or fine-tuning open-weight large language models, including supervised fine-tuning and domain adaptation., • Hands-on experience with reinforcement and alignment methods, such as RLHF, RLAIF, or direct preference optimization., • Deep experience designing and running model evaluation, including accuracy, groundedness, calibration, and safety., • Strong Python and modern deep-learning frameworks, and comfort with distributed training., • Rigor about data quality, contamination control, and honest interpretation of results., • A track record of turning training and evaluation work into models that shipped and held up., • Care for safety, bias, and trust as part of how models are built.Preferred qualifications, • Experience with the open-weight model ecosystem and parameter-efficient fine-tuning., • Experience with reward modeling, preference-data design, and large-scale distributed training., • Familiarity with retrieval-augmented generation, grounding, and knowledge graphs., • Familiarity with neuro-symbolic methods and how structure supports reasoning quality., • Experience building evaluation frameworks and quality gates for production models., • Publications, patents, or shipped systems in model training, alignment, or evaluation., • Domain understanding of knowledge-intensive or professional communities.How you work, • You define what good means before you start training, and you measure against it honestly., • You are honest about results, including negative ones, and never make a model sound better than the evidence supports., • You balance capability against cost, latency, safety, and what production can bear., • You treat trust, calibration, and safety as first-order, not afterthoughts., • You make your work reproducible so others can trust and build on it., • You lead by raising the standard and mentoring the people around you.Skills & Competencies, • Open-weight model training, fine-tuning, and domain adaptation., • Reinforcement and alignment, including RLHF, RLAIF, and preference optimization., • Reward modeling and preference-data design., • Evaluation design for accuracy, groundedness, calibration, safety, and robustness., • Training and evaluation data quality, including contamination control., • Distributed and efficient training at scale., • Grounding and retrieval-aware training and evaluation., • Quality gates, regression testing, and production model monitoring., • Reproducible, well-documented experimentation., • Technical leadership and mentorship in model quality.Services & Tools Experience, • PyTorch and the modern deep-learning stack., • Training and fine-tuning tooling for open-weight models, including parameter-efficient methods and reinforcement or preference-optimization libraries., • Distributed training frameworks (for example FSDP, DeepSpeed, or Megatron-class systems)., • Evaluation and benchmarking frameworks, plus experiment tracking and versioning., • Data-curation, deduplication, and quality tooling for training and evaluation sets., • Retrieval, embeddings, and graph access for grounded training and evaluation., • GPU and accelerator environments and cloud platforms (AWS, GCP, or Azure)., • Python as the primary language, plus solid software engineering practice., • Feeding trained models into the COGENT architecture and the MINERVA platform (trained on the job).Prior Experience & Background, • Prior work training, fine-tuning, or aligning large language models at a software, AI, or research organization., • Experience owning model evaluation and quality for systems that shipped., • A background that bridges research rigor and production engineering., • Experience with reinforcement or preference methods on real models., • Experience setting standards or mentoring others in model training or evaluation is a plus.Cross-functional partners You work most closely with Research, Foundational Model Research, Neuro-Symbolic AI, Applied AI, ML Infrastructure, and Data and Knowledge Engineering. You define, lead, and perform the training, reinforcement, and evaluation that produce the models behind the COGENT architecture and the MINERVA platform. How we hire We review every application, and we encourage you to apply even if you do not match every line above. Research shows that talented people, especially those from underrepresented communities, often hold back when they do not meet every qualification. If that is the only thing holding you back, apply anyway. Sapience AI is an equal opportunity employer. We are committed to a workplace where everyone, regardless of background, has a voice in building what comes next. Compensation Base Salary: $204,000 - $216,000 + early stage equity Generous health and wellness benefits Sapience AI is an equal opportunity employer. We do not discriminate on the basis of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. If you need an accommodation to complete our application process, let your recruiter know.