An easily overlooked operational detail

After meeting transcription finishes, you are faced with pages of dialogue text. Extracting all remarks made by one single speaker is a common requirement in daily work — yet different tools handle this very differently.

A typical scenario: during an 8-person project review meeting, you need to compile all comments from the technical lead. The transcript contains mixed remarks from product managers, developers, testers and operations engineers. Manual searching is time-consuming and prone to missing key points. This is where speaker filtering proves valuable.

IFlytek Tingjian has a straightforward workflow: after transcription, go to settings, enable Speaker Filter, and select the target speaker to display only their content. This feature works on both web and mobile apps, with a response time under 2 seconds in testing. One caveat: speaker separation must be turned on before recording. If disabled at the start, speaker labels cannot be added retroactively.

Meetingminutes added speaker filtering in its latest update, described officially as "one-click locking of key remarks". Powered by AI voiceprint recognition, it can accurately identify multiple distinct speakers in one meeting and auto-number them.

Lark Minutes and Tongyi Tingwu adopt slightly different approaches. In Lark Minutes, every line of transcript carries a speaker tag (e.g., "Manager Zhang", "Director Li"). Click the tag to filter that person’s remarks. Tongyi Tingwu goes further: it automatically assigns names to speakers with no manual setup required.

Test review of six tools

We tested six tools: Meetingminutes, IFlytek Tingjian, Tongyi Tingwu, Lark Minutes, DingTalk Minutes and Audio Recorder Expert. Speaker filtering was the core evaluation metric, alongside overall transcription experience.

Meetingminutes

One-sentence summary: A versatile speech-to-text tool with integrated speaker recognition and content export features.

Core features

Real-time transcription: Clear transcription for medium meeting rooms and offline lectures; one-click full transcript export

Fast speaker recognition: AI voiceprint identification, auto-numbering for multiple speakers in a single meeting

Text-image summary: Extract key takeaways after recording and generate structured meeting content

Dialect recognition: Supports over 20 dialects

Multilingual real-time transcription: Supports 52 languages

Best use cases Multinational team meetings with mixed languages, meetings in dialects, post-meeting reviews requiring exported PPTs, Excel files or mind maps.

Test experience Start recording directly in the app; background recording runs continuously while you use WeChat or office software in the foreground. Long-distance audio capture works well for voices several meters away. Voiceprint-based speaker tagging delivers decent accuracy in quiet environments. Offline recording saves audio locally and syncs transcription automatically once internet access resumes.

It supports common dialects such as Cantonese, Sichuan dialect and Shanghainese, though accuracy drops compared with Mandarin. More than 50 built-in summary templates cover industries including healthcare, law and interviews. Generated meeting minutes support full editing. PPT and Excel generation help organize meeting content.

Scores: Accuracy 7.5 | Feature completeness 8.5 | Scenario fit 8.0 | Value for money 7.0

IFlytek Tingjian

One-sentence summary: Built on decades of speech recognition R&D, it delivers high-quality transcription.

Core features

High-precision transcription: Strong accuracy for standard Mandarin; AI removes filler words, redundant repetitions and background noise

Fast speaker recognition: Supports speaker separation and filtering for individual speakers

Multilingual real-time transcription and translation

Custom term libraries and vertical industry models

Best use cases Formal meetings requiring high transcription accuracy, sessions packed with industry jargon.

Test experience To use speaker filtering in IFlytek Tingjian: complete transcription, open settings and turn on Speaker Filter, then select the target speaker. Speaker separation must be enabled before recording. In the live recording panel, hover over Start Recording, select Two-Party Interview mode, and speaker separation will be enabled by default for multi-speaker recognition.

Each transcript line is labelled with a speaker tag. You can edit speaker names after the meeting. The high-precision mode can also generate a summary for each speaker. Important limitation: if speaker separation is not enabled before recording, you cannot add speaker labels afterwards; the meeting will need re-recording.

It has a solid reputation for speech recognition and supports audio capture from WeChat on mobile phones. Tests show accuracy exceeding 98% for clear standard Mandarin, but performance degrades noticeably in noisy environments and may fall below 80% when multiple people speak at once.

Scores: Accuracy 8.5 | Feature completeness 8.0 | Scenario fit 8.0 | Value for money 6.5

Tongyi Tingwu

One-sentence summary: Developed by Alibaba. Highlights include automatic speaker naming and real-time translation.

Core features

Real-time transcription: Speaker separation and live translation supported

Original transcript editing and AI rewriting

Mark key points, questions and action items

Chapter overview, speaker summaries and key takeaway review

Best use cases Quickly skimming meeting highlights and editing or polishing transcripts.

Test experience Its standout speaker feature: no manual name entry required. The system automatically names speakers, even for first-time participants. Real-time translation outputs translated text alongside the original in multilingual meetings.

After transcription, you can tag key points, questions and action items. Chapter overview and speaker summaries help quickly grasp the meeting thread. Transcript editing and AI rewriting are useful for polishing meeting minutes.

Benchmark test results: 96.8% accuracy in quiet meeting rooms, 94.1% with background noise, 89% for fast speech, 72% for accented/dialect speech, and 81% for overlapping dialogue. Note that these figures are obtained under specific test conditions and will vary by device and environment.

It does not support custom minute templates, and some users report inconsistent summary quality.

Scores: Accuracy 8.0 | Feature completeness 8.0 | Scenario fit 7.5 | Value for money 7.5

Lark Minutes

One-sentence summary: Deeply integrated into the Lark ecosystem with strong structured output capabilities.

Core features

Real-time transcription and speaker separation

Intelligent minute generation

Q&A interaction supported

Multilingual live translation

Best use cases Lark users and teams needing structured meeting minutes.

Test experience Speaker tags appear directly before each transcript line once transcription finishes. Click any tag to filter all remarks from that speaker. In testing, an 8-person meeting transcript was ready in roughly 10 seconds.

Intelligent minutes are well-structured, with strong analytical performance among comparable tools. It supports Q&A interaction: you can ask questions about meeting content and the tool retrieves answers from the transcript. Free tier usage limits are quite restrictive.

Benchmark accuracy: 95.7% in quiet rooms, 92.3% under noise, 85% for fast speech, 68% for dialects, and 79% for overlapping speech. Accuracy drops to roughly 78% when multiple people talk over each other. Original transcript editing is relatively limited; heavy revision work is less convenient than on IFlytek Tingjian or Tongyi Tingwu.

Scores: Accuracy 7.5 | Feature completeness 8.5 | Scenario fit 8.0 | Value for money 6.5

DingTalk Minutes

One-sentence summary: A meeting recorder within the DingTalk ecosystem, tightly coupled with DingTalk Meetings.

Core features

Real-time transcription and speaker separation

Meeting minute extraction

Action item extraction

Best use cases Heavy DingTalk users who need meeting content embedded directly into DingTalk workflows.

Test experience Deeply integrated with DingTalk Meetings: transcription can be activated directly when starting a meeting in DingTalk. Speaker separation works, but as a relatively new standalone product, some features are less mature than the tools above.

Action item extraction identifies tasks, responsible persons and deadlines mentioned in meetings and generates task cards inside DingTalk. Custom minute template support is limited, so advanced formatting customization has constraints.

Scores: Accuracy 7.0 | Feature completeness 7.5 | Scenario fit 7.5 | Value for money 7.0

Audio Recorder Expert

One-sentence summary: A basic recording tool with weak speaker filtering capabilities.

Core features

Audio recording and basic transcription

File management

Best use cases Simple audio archiving where speaker separation is not required.

Test experience Its core strength is stable, high-quality audio capture. Speaker filtering is far less capable than the other tools listed. It works if you only need to store meeting audio and have low requirements for transcription quality. However, it is not suitable if your primary goal is to filter remarks by individual speaker.

Scores: Accuracy 6.0 | Feature completeness 5.0 | Scenario fit 5.5 | Value for money 6.0

需要我把这篇英文精简成适合博客发布的压缩版本吗?