Sākshāt Goyal
Sākshāt GoyalProduct Designer
LinkedInResume

Creating an AI-driven research architecture for reliability and novel exploration.

As AI products multiplied around automation and content generation, HBS wanted to explore a more meaningful opportunity for AI in strategy among executives. My goal was to identify potential users and problems worth pursuing. The work resulted in:

  1. A defined problem set, frameworks to interpret and prioritize them, and follow-on product discovery.
  2. Human-AI interaction principles used to fine-tune model responses across HBS projects for business executives.
  3. A research architecture to quantify qualitative responses at scale for future research projects.
Stakeholder
HBS AI Institute (previously D^3 Institute)
Skills
Research StrategyProduct DiscoveryQualitative Data Analysis
Year
2025 (4 weeks)
Role
Designer + Researcher

The team

The HBS AI Institute engaged us to lead the research and design work, with the four team members below serving as project advisors.

  • Tanya FlintHead of AI
  • Yogesh KumarSr. Data Engineer | Project Advisor
  • Jean-Luc JacksonSr. Data Scientist | Project Advisor
  • Rem KoningAssistant Professor | Project Advisor

Problem:

There was no defined industry, executive role, or problem space. “Executives” could mean anyone from a CEO or chair to a director or team lead. “Strategy” could include almost every consequential decision they made.

Without clear boundaries, almost any polished answer could appear complete, even when important situations had been missed.

Feeling untethered without a well-defined audience, industry, or problem space. [AI-generated]

Approach:

To avoid narrowing the opportunity around our initial assumptions, I organized the work around three questions:

  1. How executives formulate strategy,
  2. What problems they encounter while doing so, and
  3. What role AI currently plays in their decisions.

I gathered credible resources and papers to examine these questions broadly across roles, industries, and types of strategy. To handle research at scale, I built an AI-assisted research architecture. The true challenge was reflecting on what we mean when using common terms like “insight,” “synthesis,” and “analysis,” then codifying those meanings for an LLM.

Data collection
Extraction
Classification
Inquiry routing

Maintaining standards for AI outputs:

Reasoning models gave reliable answers but missed hidden meanings. Creativity models could go beyond basic understanding, but they weren’t consistent in applying logic across different sets.

Prompt engineering let us combine GPT-4.5’s creative and analytical range with o3’s consistency. I kept the goal fixed while testing several versions of the AI pipeline.

I approached prompt design as a collaboration with a distributed team of designers and researchers, each with their own opinions and implicit biases. This helped balance situational judgment with consistency.

Mapping current models’ strengths and weaknesses against our target performance.
Models compared by consistency and interpretive depth to find the strongest balance between reliable output and meaningful interpretation.

Designing situational axes:

Despite several attempts at sorting, the clusters remained too broad for meaningful collaboration or productive discussion.

I considered subdividing the clusters by industry, but that risked reinforcing industry stereotypes. Activities within one industry can differ substantially; pharmaceutical marketing and legal operations within a content-generation team, for example, are distinct functions.

I mapped each activity along shared situational axes, including regulation, modularity, timing, value horizon, knowledge transfer, and market spread.

This made the collections mutually exclusive and collectively exhaustive, allowing every module to be mapped.

Each module tagged by the conditions influencing the activity.

Once the clusters were manageable, I used contrastive typology: scoring evidence along bipolar situational dimensions and comparing their intersections to identify distinct operating contexts. The dimensions produced several plausible permutations.

Collaborative sessions helped us remove factors that did not create meaningful distinctions, support the team’s objectives, or inform useful product decisions.

Four situational axes created a common language for comparing otherwise unrelated forms of strategic work.

Designing principles of human-AI interaction:

Although outside the original scope, I recognized that the data from Research Question 3 could help us define principles for AI interactions with executives during strategic discussions.

Design principles often sound good but aren’t actually useful. When they become truisms like “AI should be trustworthy,” “prioritize the user,” or “support human judgment,” they’re easy to agree with but hard to use in real design work. I wrote each principle so its opposite could also make sense—something a smart team might choose in another situation or with different goals.

This helped me develop a way to frame principles as thoughtful trade-offs, grounded in context rather than as absolute truths.

Impact:

The work continued in three ways.

  1. The seven human-AI interaction principles were used to fine-tune model responses across exploratory HBS projects for business executives. They gave teams specific behaviors to compare. For example, whether a model should present its conclusion or reasoning first, accept the user’s framing or introduce another perspective, and state uncertainty directly or leave it implicit.
  2. The research architecture was repurposed for other qualitative work.
    We used it whenever we needed to convert unstructured text into comparable measures and reviewable patterns. Two immediate examples included scoring participant-written prompts at scale for HBS’ Sana.ai prompt engineering course and analyzing open-ended surveys from Leading with AI events.
  3. The sixteen people problems initiated follow-on research and exploration across HBS teams. Instead of beginning with “executive strategy” as one enormous subject, teams could investigate specific tensions such as deployment versus governance, short-term gains versus long-term capability, or intuitive judgment versus structured analysis.
Research architecture repurposed for additional qualitative analysis across the org.

Lessons:

The hardest part was distinguishing genuine insight from plausible-looking output.

Lesson #1: AI forced me to become more precise about research.

To get precise outputs, a lot of effort went into defining and revisiting terms we often take for granted. For example: “What makes an insight insightful?” and “Which form of thematic synthesis is most applicable to a problem?”

Lesson #2: At scale, AI-assisted research can very quickly turn into a pipeline management activity.

Once hundreds of sources moved through the study, prompts, definitions, scoring rules, model choices, and review checkpoints became part of the research method.