Sully.ai Logo

Sully.ai

Applied Research Scientist

Reposted 6 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in US
180K-220K Annually
Mid level
Remote
Hiring Remotely in US
180K-220K Annually
Mid level
As an Applied Research Scientist, you will design experiments, build evaluation frameworks for LLMs, and ensure safety and reliability in production systems.
The summary above was generated by AI

About Us
๐Ÿ‘ Team from OpenAI, DeepMind, NASA, GoogleX, Tesla, and 2 physicians: 6 exits, 2 IPOs.
๐Ÿ”ฅ Our model outperforms Claude, Gemini, and GPT-4.5 on clinical benchmarks.
๐Ÿ“ˆ 400+ healthcare orgs signed in 16 months.
โšก๏ธ $25M raised from YC, Amity Ventures, Sequoia scouts, and more.
๐ŸŒŽ $1T+ market opportunity. Weโ€™re going after all of it.

About the Role
We are seeking an Applied Research Scientist to design and run rigorous experiments for LLM-based agents, with a focus on clinical and agentic reliability. This role will own the development of automated evaluation frameworks, bridging research prototypes into production systems, and partnering closely with product and engineering to ensure safety, robustness, and measurable impact.

Key Responsibilities

  • Design and run experiments to measure accuracy, robustness, and hallucination rates in LLM agents.

  • Build automated evaluation pipelines (LLM-as-judge + human review) with clinical-grade benchmarks.

  • Partner with Research Ops/IRB to design efficacy studies and align with regulatory requirements.

  • Translate research into production-ready evaluation systems, collaborating with engineering to land features 0โ†’1.

  • Develop error taxonomies, ablations, and guardrails to ensure safe and reliable agent behaviors.

Hard Requirements

  • Proven experience designing agentic processes and LLM evaluation/benchmarking frameworks.

  • Strong Python and ML background (PyTorch/TensorFlow, Hugging Face, LangChain/LlamaIndex).

  • Demonstrated ability to design rigorous experiments and translate findings into production.

  • Track record of published research or deep applied work in LLMs and agent evaluation.

  • Strong communication and technical writing skills to articulate complex findings clearly.

Nice-to-Have

  • Prior work in healthcare/clinical NLP with awareness of medical data standards.

  • Experience running IRB-aligned or clinical-grade studies.

  • Exposure to noisy/limited medical data and designing strategies to overcome constraints.

First-Month Focus

  • Audit existing evaluation approaches for clinical and agentic tasks.

  • Define initial benchmarks and build early automated pipelines.

  • Partner with engineering to land first set of CI gates for accuracy, factuality, and safety.

Success OKRs (90 Days)

  • Deliver a repeatable evaluation framework with automated pipelines in production.

  • Demonstrate measurable improvements in robustness, hallucination reduction, or safety.

  • Publish or present internal research findings that directly shape product reliability.

Culture Fit

  • Persistent, driven problem solver

  • Willing to push back on leadership to defend quality/timelines

  • Thrives in high-ambiguity, fast-paced startup environments

Why Join Sully.ai?
๐Ÿ”ฅ Shape the Future of Healthcare: Build category-defining partnerships that enable doctors to focus on saving lives.
๐Ÿ“ˆ Early-Stage Impact: Join early and play a critical role in shaping our partnership roadmap and overall company growth.
๐ŸŒŽ Remote-First Culture: Work with a talented, mission-driven team in a flexible, remote environment.
๐Ÿ’ฐ Competitive Compensation: Enjoy a competitive salary, equity, and the opportunity to make a real difference.
๐Ÿ† Solve Scalability Challenges: Tackle complex challenges in a rapidly growing company, driving impactful change in healthcare.

Sully.ai is an equal opportunity employer. In addition to EEO being the law, it is a policy that is fully consistent with our principles. All qualified applicants will receive consideration for employment without regard to status as a protected veteran or a qualified individual with a disability, or other protected status such as race, religion, color, national origin, sex, sexual orientation, gender identity, genetic information, pregnancy or age. Sully.ai prohibits any form of workplace harassment. 

Top Skills

Hugging Face
Langchain
Llamaindex
Python
PyTorch
TensorFlow

Similar Jobs

3 Days Ago
In-Office or Remote
2 Locations
190K-300K
Expert/Leader
190K-300K
Expert/Leader
Software
As a Staff Applied Research Scientist, you'll oversee deep learning research, mentor team members, and drive algorithm development for media synthesis and enhancement challenges.
Top Skills: Deep LearningPyTorchTensorFlow
40 Minutes Ago
Easy Apply
Remote
2 Locations
Easy Apply
135K-186K Annually
Senior level
135K-186K Annually
Senior level
Artificial Intelligence • Fintech • Machine Learning • Social Impact • Software
Lead and inspire a team of recruiters focused on technical hiring, own engineering recruiting strategy, and build effective processes to enhance candidate experience and drive efficiency.
40 Minutes Ago
Easy Apply
Remote
2 Locations
Easy Apply
164K-226K Annually
Senior level
164K-226K Annually
Senior level
Artificial Intelligence • Fintech • Machine Learning • Social Impact • Software
As a Senior Software Engineer, you'll lead projects improving loan origination platforms, build reliable APIs, and enhance customer experience with a focus on quality.
Top Skills: AWSDockerGithub ActionsKotlinReactRubyTypescript

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the cityโ€™s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Loweโ€™s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account