GPT-4o brings real-time voice
OpenAI's GPT-4o handles text, audio and images in a single model, enabling spoken conversations with response times close to a human's.
A historical account of this event. Usage notes are YLEM editorial analysis; check the provider for current features and prices.
What changed
GPT-4o was introduced as a model trained across text, vision and audio. The demonstrated abilities and the features rolled out to users were not identical: the announcement described staged availability.
Sources: [1] OpenAI · Hello GPT-4o
What this means in practice
A useful multimodal task combines a question with relevant material: a chart to explain, a screenshot to inspect or a spoken practice conversation. Tell the assistant what kind of help you need.
How to read the result
Ask it to separate what is visible from what it infers. When reading a chart, verify axis labels and units before relying on the conclusion.