Shipped an AI co-pilot scientists could trust, saving them hours of work

0 to 1Self-InitiatedAI PlatformShipped

CHALLENGE

BenchSci’s products helped scientists find important insights for their experiments, but scientists still had to read and synthesize the dozens (sometimes hundreds) of surfaced papers.

GenAI could help close that gap, but AI also hallucinated and most scientists didn’t trust it. Could synthesizing experiments with GenAI help scientists get to insights faster without breaking the trust the BenchSci products were built on?

WHAT I DID

I made the case for the research, then led design on the first GenAI feature. AI became an assistive layer scientists had to trigger, not an authority baked into the interface. Every claim got anchored to a citation, so trust travelled with the output. Rather than wait for full paper coverage, I shipped analysis of the top 20 papers, with that limit stated right next to the answer instead of hidden.

OUTCOME

This feature cut the time it took for scientists to get insight in the Navigator product. Scientists reported getting there faster, and quantitative measures of adoption backed it up. They used the AI summary, checked its citations, and copied it into their own work. That validated that scientists will use AI when the verification is traceable and that AI can speed up how they get to insights.

It also created the citation-and-trust pattern other teams adopted on their own, written up as a five-part guideline they designed against.

This became the company’s first GenAI feature. The AI direction fed a US$70 million Series D and three eight-figure contracts.

ROLE

Principal Designer and primary owner. No product manager was assigned, so I also selected the first use case and drove the product trade-offs.

TIMELINE

Apr to Jul 2023

TEAM

Started solo. Grew to a tiger team of five: a senior designer on detailed pattern design, engineering, a project manager, and a scientist.

DETAILED PROJECT BREAKDOWN

Background: Scientists use BenchSci to understand disease biology and validate experiments before costly drug-program decisions

e.g. Does this protein drive this disease? What does the literature actually say about this target? A wrong call is measured in years and millions of dollars.

BenchSci's products, like Navigator (below), surfaced the relevant evidence. However, it still left scientists to read and synthesize dozens (sometimes hundreds or thousands) of scientific experiments and papers.

BenchSci's Navigator product let scientists to search for a protein, disease and other disease context and surface relevant research papers.

Problem: scientists could find relevant evidence, but synthesizing it into an insight still took hours. Any AI solution risked breaking trust.

Opportunity: accelerate the insight-finding by synthesizing the evidence and presenting the most relevant parts of experiments.

Risk: AI can hallucinate. Scientists are people trained to interrogate methodology. They cannot accept a conclusion because it sounds authoritative. A pharma scientist has to show what evidence they reviewed.

Design question: Could synthesizing experiments with GenAI help scientists get to insights faster without breaking the trust the product was built on?

Nobody had asked me to answer that. BenchSci was 300+ people with 20+ enterprise contracts and a committed roadmap. Horizon scanning was part of a principal designer’s job, and I made the case for running discovery research on this question.

BenchSci was 300+ people with over 20 enterprise contracts and a roadmap.

Early research created the direction

I initiated research to test whether this deserved investment.

  • Observed tension: speed vs confidence

  • Key insight: verification had to be lightweight and local

  • Core need: traceability, in service of speed, not as an end in itself

Two directions were explored: bolt an AI chatbot on top of the existing interface, or redefine search itself around evidence and citations. Testing showed the second had far more impact.

I defined a principle: trust through visibility

The interface's job was to lower the cost of verification.

That principle guided this feature, and every AI decision that followed it.

This research fed a vision prototype that got the company and board aligned on direction and led to the formation of a tiger team.

A tiger team formed, and I prioritized the first feature by mapping user value to feasibility

My role: principal designer and primary owner

  • selecting the first use case

  • driving the core interaction and product tradeoffs

  • defining the pattern we wanted to scale

Team

  • I partnered with a senior designer on detailed pattern design

  • engineering: feasibility + implementation

  • project manager: coordination

  • scientist: domain input / QA

Prioritization:

  • I mapped where scientists lost the most time, reading and cross-referencing dozens of papers one by one, and picked the step GenAI could realistically shorten in the time we had.

  • Impact was cross-referenced with feasibility via speaking with the engineers.

Feature walkthrough:

Decision: frame AI as an assistive layer, not an authority

Engineering wanted to avoid muddying the existing evidence interface.

That was a constraint that I used: Instead of blending AI output into the interface, I made it a distinct, explicitly triggered layer. Scientists asked for a summary. It appeared as something clearly generated, clearly separate, clearly optional.

Decision: anchor every claim to an inline citation

Hover previews showed the source without leaving the page. Click-through allowed deeper inspection for anyone who wanted it.

Copying a summary copied its citations too, so trust travelled with the output, into the slide deck or the email or the report where the decision actually got made.

Citations truncate after three, with a “+X more” link that expands on click. Showing all citations inline made the text overwhelming and hard to read. Scientists read this under time pressure, and a sentence broken up by six bracketed references is a sentence nobody finishes. Truncating balances transparency with readability, and the full list is one click away for anyone who wants it.

Decision: expose scope limits without killing confidence

V1 could only analyse around 20 papers. Users wanted comprehensiveness, and engineering constraints made full coverage infeasible in the first release.

My decision: ship narrower scope, but surface the limitation near the output.

The scientist can see exactly what was and was not read, in the place where they are deciding whether to trust the output.

Generation fails sometimes. If the server was receiving too many requests, it would not work. We needed to build in a count down to inform the user when they could try again, rather than having them hit the same wall repeatedly.

What I rejected

The summary inline in the publications list. From the UX side, long summaries pushed the publications down the page, cutting into the evidence scientists came for. From the technical side, engineers didn’t want to muddy the existing app. Two constraints, same direction.

A warning banner above the output. I felt this drew too much attention to our limitations before the feature had shown any value. The goal was confidence in what we can do, not a warning about what we can’t.

Every citation shown inline. I wanted maximum traceability and tested it directly. It made the text overwhelming and hard to read, which costs more trust than it buys. Truncating at three kept the transparency and gave back the reading line.

Outcome: the time scientists got to insights dropped from hours to minutes

Scientists went from getting insights via reading and cross-referencing (estimated 15min to 2 hours of work*) to getting insights via the AI summarization feature (minutes of work).

Qualitative signals: Scientists reported being able to understand a relationship between entities in the Navigator product much faster and easier than before. They were excited by the new feature and pushed for its expansion.

Analytics showed the feature becoming part of the workflow:

  • Clicking Summarize: whether a skeptical audience opts in at all

  • Hovering and clicking a citation: whether verification is actually happening

  • Copying to clipboard: the hardest to earn, it puts the output into the scientist's own record, where they stake their credibility on it

*A rough model for time saving shift

  • Before: 20 to 50 relevant papers once narrowed to context, 30 seconds to four minutes each, then synthesise across them yourself. 10 minutes to two hours before any decision.

  • After. A synthesis of the top 20 papers. Every claim anchored to an inline citation, verifiable on hover, copyable with its citations attached. Minutes, not hours.

  • This before/after is a model built from workflow timings, not a measured result.

Company home page showing new direction and the first GenAI feature launched.

Impact: This work expanded and shaped the broader AI direction

  • The company backed this work with people post-launch: expanded to risk summaries, given a dedicated PM

  • Other teams adopted the citation design pattern

  • The flagship co-pilot (see below) launched a few months later in 2024, with an evolved citation pattern.

The trust principle and work became the foundation for the co-scientist's evolved citation pattern.

A later version of the co-scientist, continuing to use the pattern.

What I would do differently: quantitative measures of time-to-insight

I chose adoption signals over a direct time-to-insight quantitative measurement. Instrumenting the actual goal would have taken time we didn't have for a first 0-to-1 release, and there was no existing baseline to measure against. That was the right call, it let us ship fast and validate the direction. But adoption only tells you scientists used the loop, clicking, checking, copying, not that they got to insight faster.

Running it again, I'd add a coarse quantitative measure of time-to-insight, to triangulate against the qualitative and behavioural signals I already had. Knowing the loop completes is worth less than knowing whether the thing I set out to change, changed.

What happened next

By 2025 the product had moved to a multi-agent architecture. The principle and citation pattern needed to evolve so that scientists could verify how the system was working under the hood. You can read about that in this case study.