Return

Designing trust signals for humans and AI

Stack Overflow · 2026

Role
Design, research & strategy
Team
PM & tech lead
Timeline
3 months · 2026

Context

Stack Internal helps companies manage internal knowledge. As the product evolved towards automated knowledge management system, knowledge needed to be accurate, up to date and safe for people and AI to act on. However, manual reviews could not keep pace, with 80% of content unreviewed for over a year.

I led the design, research and strategy for a scoring system that made reviews more efficient and helped people and AI identify knowledge they could trust and safely act on.

Stack Internal dashboard for managing and reviewing internal knowledge

The problem

Stack Internal imported knowledge from Confluence, Google Drive, Slack and other sources. Some was useful, while some was outdated, duplicated or conflicted with another source. Content reuse had fallen by 40%, and neither people nor AI could easily tell which knowledge was safe to act on.

Impact

What I did

Understanding what experts needed to trust knowledge

Interviews and concept testing showed that trust depended on content quality, freshness, source credibility and author reputation. Experts valued human content for its depth and lived experience. AI-generated content needed clear evidence and credible sources to earn the same trust.

Experts also rejected scores they could not question. They wanted to understand how each judgement was made.

A remote interview session with a knowledge expert
Expert interviews. Understanding how experts assessed human and AI-generated knowledge.
The distrust path: opening a doc and trusting it, then doubting it through duplication, recency, accuracy and provenance until trust collapses and they ask a person
The distrust path. Where confidence in knowledge breaks down.
Concept testing results: 90% of participants would review the content in more detail before deciding, 60% would improve it first, and 30% felt the score alone was not enough to decide
Concept testing. Testing whether experts could understand and act on the score.

Making every score explainable

Every score had to show how it was generated. I used explainable signals to help experts assess the reasoning, identify uncertainty and prioritise reviews.

Scoring widget. Recommended next steps, flagged content requiring expert review and explained how each score was generated.
Ingestion review. Showed experts what to review, why it was flagged and what needed to change in imported knowledge.
Scoring variation. Yellow marked knowledge that needed a fresh review. Red marked content that was unsafe to act on.

Learning

People trusted the score when they could see how it was reached.