ChatGPT Now Supports Audio Files | Transcription and Summarization Now Available
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
Have you ever spent hours listening back to a meeting recording just to put together the minutes? Now you can hand ChatGPT an audio file and have it transcribed and summarized for you. This article walks through what the new feature does, where it fits in, and what to watch out for.
US-based OpenAI announced the audio file upload feature in its release notes on October 6, 2026. ITmedia AI+ also covered the news on October 9.
There are three main things you can do: transcribe a recording (converting spoken words into text), summarize the content, and ask questions about it.
Until now, ChatGPT had limited ability to read audio directly. You had to use an external transcription tool to convert audio to text first, then paste it in. This new support eliminates that extra step.
Note that this is separate from the existing "Record" feature. Record is a function in the Mac app for capturing audio on the spot, whereas this new feature lets you upload files you already have.
This feature is for paid plans. According to a detailed article, the official help documentation states it is available "with a paid ChatGPT subscription and workspace." Free plans are not currently eligible.
The article also notes that availability may vary depending on region, app version, and the model selected.
Supported formats are WAV, MP3, OGG, PCM, FLAC, AAC, and M4A. Audio-only WebM and MP4 files are also accepted.
The limit is 512 MB per file. No maximum playback duration is specified. However, very long recordings may be processed in segments or time out (where processing stops due to exceeding the time limit).
WebM and MP4 files that are identified as video are not accepted. If you want to transcribe a video file, you will need to extract the audio track first.
Note that some reports contain descriptions of supported formats and file deletion handling that differ from what is presented here. For the latest details, please check OpenAI's official help documentation.
The steps are simple. Open a new chat, attach your audio file, then write a specific description of what you want and send it.
If you just say "summarize this," you tend to get a vague response. Prompts like the following will get you more actionable results:
The detailed article also recommends adding a note like "do not fill in names or numbers by guessing" — this prevents the model from making up details it could not hear clearly.
You can also ask follow-up questions within the same chat. For example, you could say "based on the decisions from earlier, draft a follow-up email," and it will generate the text for you.
Consider a sales rep at a small company. They have three client meetings a week and record each one on their phone. Every evening, listening back to recordings and writing up notes was taking at least an hour.
With this feature, they can simply upload the recording and ask: "List the customer's requests and next actions." Verification is still needed, but the time spent listening back can be significantly reduced.
What about a university student? They might record a 90-minute lecture and want to review only the key points before an exam. By uploading the recording and asking "list 10 important terms with explanations," they can create study notes.
For freelance writers and YouTubers, the feature is useful for organizing interview recordings. Getting a rough sense of who said what can cut down the time spent hunting for specific quotes. That said, always verify any quotes you plan to use against the original audio.
ChatGPT is not the only AI that can handle audio files. According to reports from September 2025, Google's Gemini accepts MP3, M4A, and WAV audio files — up to 10 files at once, with a combined limit of 10 minutes. It reportedly supports transcription, summarization, extraction of action items, and speaker identification.
These limits are from that time and may have changed since. Check Gemini's official help documentation before using it.
ChatGPT's advantage is its 512 MB per file capacity, which means longer meeting recordings may fit in a single upload.
On the other hand, dedicated transcription apps like Notta and Otter are purpose-built for meeting recordings. They generally offer features tailored for meetings, such as speaker separation and timestamped transcripts. For situations where a verbatim record (word-for-word documentation) is required, a dedicated app is likely the safer choice.
A practical way to divide usage: use ChatGPT when you want a rough grasp of the content to move on to the next task, and use a dedicated app when you need an accurate, reliable record.
For readers in Japan, the accuracy of Japanese transcription is a natural concern. The ITmedia AI+ article contains no specific mention of availability in Japan or Japanese language accuracy.
OpenAI has noted that accuracy in understanding and transcribing audio may vary by language. Speaker identification can also be incorrect.
Other reports describe it as working best in English, and note that accuracy drops with background noise, accents, or overlapping speech. Japanese business meetings often involve fast speech, technical jargon, and multiple people talking at once, so always verify the output with your own ears.
Privacy is also worth considering. According to the detailed article, there is no indication that uploaded audio is deleted after processing. Retention depends on where the file is stored and the workspace's data retention settings.
For personal plans, use of data for training depends on your data settings. Business plans are treated as not using customer data for training.
Before uploading recordings of client meetings or confidential internal discussions, check your company's policies. It is also important to inform participants that the session is being recorded and that AI tools will be used, and to obtain their consent in advance.
Not at this time. It requires a paid plan and its associated workspace.
No maximum playback duration is specified. The file size limit is 512 MB per file. Very long recordings may be processed in segments or time out.
Files identified as video are not accepted. Extract the audio track and convert it to MP3 or another supported format before uploading.
This is not recommended. The output may contain errors or mix up who said what. Always verify numbers, dates, and proper nouns against the original recording.
It is recommended that you inform all participants that the session is being recorded and that AI will be used to create notes, and obtain their consent beforehand. Rules vary by country and company, so be sure to check.
Start by uploading a short recording you have on hand and see how well the summary turns out.
This article is a cross-post from AI Friends.