Data Scientist 3055-1
Data Science
Remote
Posted on Aug 19, 2026
Our client is currently seeking a Data Scientist 3055-1
Job Description: The Role
We are looking for an Applied AI Engineer / Scientist to train, build, evaluate, and continuously improve our production pipelines.
You will work at the intersection of software engineering, model training, LLM systems, evaluation, and deep healthcare workflow understanding. Your role is to put state-of-the-art model capability into production in reliable, auditable and maintainable pipelines, and continuously improve from real-world failures.
You will be embedded in hard healthcare problems — clinical documentation integrity, autonomous medical coding, denial prevention, appeals, revenue cycle workflows, and payer logic — and will own the loop from problem framing, model training, agents, tools, delivery, evaluation and improvement.
The ideal candidate is a curious applied scientist with an engineer mindset: rigorous about measurement, comfortable with ambiguity, excited by messy real-world data, and motivated by making things real. You want to close the gap between impressive demos and dependable production systems.
What You’ll Do
• Support a production AI system end to end from research to deployment for a product.
• Design, build, and iterate on the production AI systems for real-world healthcare workflows that demand extreme accuracy and auditability with complex label spaces.
• Build evaluation suites that measure system performance across generative, extraction, retrieval, and classification tasks, including multi-label and multi-class classification problems, regressions, edge cases, safety, reliability, provenance quality, and business impact.
• Analyze human-in-the-loop feedback data to continuously raise the bar of model performance
• Work with clinical, coding, product, and operations experts to translate domain workflows into scoped production AI systems
• Build feedback loops from expert review, production logs, model outputs, and benchmark runs
• Research and prototype new models and capabilities from the ground up, run evaluations and move the best into production
• Partner with research scientists and ML engineers on model selection, supervised fine-tuning, reward modeling, distillation, synthetic data generation, or post-training experiments.
• Ensure outputs are clinically useful, explainable, and auditable, with clear evidence, source provenance, and decision rationale.
• Help define what “good” looks like for AI systems completing complex tasks end-to-end.
You May Be a Good Fit If You
• Have 4+ years of software engineering, ML engineering, research engineering, or applied AI experience.
• Are highly proficient in Python and comfortable building production systems with APIs, structured data, async workflows, testing, logging, and observability.
• Have experience training or fine-tuning neural networks for classification or generative tasks
• Have experience framing messy real-world workflows as structured prediction problems, including multi-label classification, multi-class classification, ranking, extraction, and decisioning
tasks.
• Have hands-on experience with LLM applications, agent frameworks, tool calling, retrieval/RAG, structured outputs, prompt design, or model evaluation.
• Have built or operated evaluation systems, benchmarks, annotation workflows, experiment tracking, or regression testing for AI systems.
• Think in terms of systems and user outcomes, not just model metrics.
• Enjoy debugging messy real-world failures and turning them into durable product and model improvements.
• Are comfortable working directly with domain experts and translating ambiguous workflows into concrete technical specs.
• Can move quickly in loosely defined environments while maintaining high standards for correctness, reliability, and safety.
• Want to work in the layer that turns model potential into systems that actually work for users.
Strong Signals
We are especially excited by candidates who have experience with one or more of:
• Experience building classifiers or evaluation systems for ambiguous, imbalanced, high-stakes domains, including multi-label and multi-class classification, hierarchical labels, threshold
tuning, calibration, and precision/recall tradeoffs.
• Production LLM systems, agentic workflows, or tool-using AI systems.
• Evaluation frameworks, LLM judges, expert-labeled benchmarks, or human-in-the-loop review systems.
• Fine-tuning, supervised fine-tuning, reinforcement learning, reward modeling, distillation, or synthetic data generation.
• Retrieval systems, search, ranking, embeddings, reranking, or long-context information retrieval.
• Healthcare, clinical documentation, medical coding, claims, denials, payer policy, or revenue cycle management.
• Building systems over large-scale unstructured or semi-structured data.
• Designing AI systems that require explainability, auditability, compliance, or safety review.
• Working in fast-moving startup or research-to-production environments.
What Makes This Role Different
Most AI roles are either too research-heavy or too product-light. This role sits in the middle.
You will not only write prompts or run experiments. You will support the core AI system that powers a live product, using both trained models and agents to deliver value to users, and continuously learning from their feedback.
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively “Judge”) to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.