Newcode.ai Logo

Newcode.ai

Head of Evaluations (Legal AI Benchmarking)

Posted 18 Days Ago
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Lead design and scaling of legal AI evaluation frameworks and benchmarks, build and maintain datasets, audit and score AI legal outputs, define metrics focused on reasoning and citation accuracy, and collaborate with engineering to translate legal errors into model fine-tuning feedback.
The summary above was generated by AI
Who are we?

At Newcode.ai, we're transforming how law firms and legal professionals harness AI for real-world impact. As part of our collaborative, high-growth team, you'll have the rare opportunity to work side-by-side with visionary founders at the bleeding edge of AI and legal innovation — shaping not just our product, but the future of legal work itself.

Note: We believe in being transparent about what it's like to work at Newcode. As a fast-growing startup, we're building and evolving every day. That means not every process, playbook, or framework is already in place, and priorities can shift quickly.

The people who thrive here are comfortable with ambiguity, take ownership, and don't wait for perfect direction. They are resourceful, proactive, and able to "figure it out"—solving problems, creating structure where needed, and helping build the company as they go. If you are good with this then, great! Keep reading to learn more.

Position Overview 

We are seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform. You will ensure our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate. 

Key Responsibilities 

  • Design Legal Benchmarks for: Contract Drafting, Information Extraction, Legal Research, and Contract Review 
  • Build, source and maintain relevant datasets 
  • Audit AI Output: Review and score complex AI-generated legal text, contract analyses, and statutory interpretations for accuracy and precision and lay out a strategy.  
  • Define Evaluation Metrics: Establish clear criteria for grading model performance, specifically focusing on logical reasoning, citation accuracy, and the model's ability to safely abstain from answering. 
  • Collaborate with Engineering: Partner directly with Engineering to translate legal errors into actionable technical feedback for model fine-tuning. 

Requirements
  • PhD or Masters in statistics, mathematics, machine learning or equivalent 
  • Analytical Skills: Proven ability to break down complex statutory frameworks and case law into structured, logical data points. 
  • Tech-Savviness: python, panda, numpy, jupiter notebooks and similar statistical models 

Similar Jobs

2 Days Ago
Remote or Hybrid
Mid level
Mid level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Provide advanced technical support for Rain's payments platform: troubleshoot APIs and integrations, read logs and run SQL queries, own incidents in 24x7 coverage, communicate with partners and internal teams, produce documentation and troubleshooting guides, and feed product and engineering insights to improve systems and roll out features.
Top Skills: Rest ApisScriptingSQL
2 Days Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Drive new business sales for a SaaS license model by prospecting, account and territory planning, engaging C-suite stakeholders, orchestrating cross-functional teams, advising on IT and AI integration, negotiating deals, and achieving sales targets.
Top Skills: AISaaSServicenow
2 Days Ago
In-Office or Remote
2 Locations
150K-250K Annually
Mid level
150K-250K Annually
Mid level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
The role involves researching and developing large language models (LLMs) with a focus on transformer architecture, data curation, distributed training, and optimization. Responsibilities include conducting experiments, collaborating with teams, and staying updated on deep learning advancements.
Top Skills: Distributed ComputingLarge Language ModelsPythonPyTorchTransformer Architectures

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account