PROMPT CANVAS CANVAS
# Role & Context
You are a Senior AI Infrastructure Engineer specializing in Speech-to-Text and NLP Pipelines. Your task is to design an automated pipeline that ingests meeting audio, performs speaker diarization (identifying who is speaking), transcribes the audio, and finally utilizes an LLM to generate a speaker-aware meeting summary.

# Input Data
- Primary Goal: {{primary_goal}}
- Desired Transcription Format: {{transcription_format}}

## Step-by-Step Instructions
1. Analyze the `primary_goal` to determine the specific downstream tasks required (e.g., action item extraction, sentiment analysis, CRM logging).
2. Within a dedicated thinking block, architect the pipeline flow:
   - **Phase 1: Audio Preprocessing & Diarization** (e.g., using open-source models like PyAnnote or a managed service).
   - **Phase 2: Transcription** (e.g., using OpenAI's Whisper model on segmented audio chunks).
   - **Phase 3: Merging** (Combining transcriptions with speaker IDs based on timestamps to match the `transcription_format`).
   - **Phase 4: LLM Synthesis** (Passing the diarized transcript to an LLM like GPT-4o for summarization).
3. Provide the overarching system architecture.
4. Write the core integration code (in Python) that glues the diarization output with the Whisper transcription and formatting logic.

## Constraints
- Ensure the code gracefully handles overlapping speech segments if the chosen diarization method supports it.
- Provide clear instructions on how to structure the final prompt sent to the LLM to maximize the utility of the speaker labels.

Role & Context

You are a Senior AI Infrastructure Engineer specializing in Speech-to-Text and NLP Pipelines. Your task is to design an automated pipeline that ingests meeting audio, performs speaker diarization (identifying who is speaking), transcribes the audio, and finally utilizes an LLM to generate a speaker-aware meeting summary.

Input Data

  • Primary Goal: {{primary_goal}}
  • Desired Transcription Format: {{transcription_format}}

Step-by-Step Instructions

  1. Analyze the primary_goal to determine the specific downstream tasks required (e.g., action item extraction, sentiment analysis, CRM logging).
  2. Within a dedicated thinking block, architect the pipeline flow:
    • Phase 1: Audio Preprocessing & Diarization (e.g., using open-source models like PyAnnote or a managed service).
    • Phase 2: Transcription (e.g., using OpenAI's Whisper model on segmented audio chunks).
    • Phase 3: Merging (Combining transcriptions with speaker IDs based on timestamps to match the transcription_format).
    • Phase 4: LLM Synthesis (Passing the diarized transcript to an LLM like GPT-4o for summarization).
  3. Provide the overarching system architecture.
  4. Write the core integration code (in Python) that glues the diarization output with the Whisper transcription and formatting logic.

Constraints

  • Ensure the code gracefully handles overlapping speech segments if the chosen diarization method supports it.
  • Provide clear instructions on how to structure the final prompt sent to the LLM to maximize the utility of the speaker labels.

Speaker-Aware Meeting Intelligence Pipeline

Productivity 2026-08-08T00:00:00Z
đŸ‘ī¸ 0

Key Capabilities

  • Automated Audio Diarization Strategy: Build out logic that flawlessly maps transcribed text segments to distinct speaker IDs, solving the "who said what" problem in multi-participant recordings.
  • Optimized Prompt Chains for Summarization: Deploy specialized NLP prompt templates to extract decisions, sentiment, and follow-up tasks on a per-speaker basis.
  • End-to-End Pipeline Architecture: Receive step-by-step guidance on connecting audio ingestion, diarization engines (like PyAnnote), and OpenAI LLMs into a seamless automated workflow.

Underlying Mechanism

This prompt utilizes a Markdown-structured Chain of Thought (CoT) to enforce a sequential processing pipeline. By cleanly separating the audio processing phase (diarization and transcription) from the synthesis phase (LLM summarization), the AI generates robust, modular code that correctly passes speaker-labeled transcripts into subsequent generative models without context loss.

Ideal Use Cases

  • Engineering internal tooling to summarize Zoom, Teams, or Google Meet recordings automatically.
  • Developing B2B SaaS products focused on conversation intelligence (e.g., sales call analysis).
  • Automating interview transcriptions for journalism, HR, or qualitative research.

When NOT to Use (Anti-Patterns)

  • Processing single-speaker audio like podcasts or monologues where diarization adds unnecessary overhead.
  • When building strict real-time applications where offline diarization latency (e.g., PyAnnote) is unacceptable.

Pro Tip

Use the {{transcription_format}} variable to specify exactly how you want the diarized output structured (e.g., [Speaker A]: Text) so the pipeline's downstream summarization prompts can parse it flawlessly without hallucinating speaker identities.

Source Reference: Build a Speaker-Aware Meeting Intelligence Pipeline with Audio Diarization

📋 How to Use This Prompt

1

Copy or Save to Vault

Click Copy Prompt for quick access, or hit the ⭐ star button to save to your Vault to edit templates, customize values, and auto-fill variables.

2

Fill in Variables

Replace double-bracket placeholders {{variable}} with your own values, context, or specific inputs.

3

Run in AI Model

Paste directly into ChatGPT, Claude, DeepSeek, or Gemini for structured, high-accuracy outputs.