Apple model combines vision understanding and image generation »

Apple researchers have published a study on a new model "Manzano" that represents improved quality and performance compared to current versions.

From Marcus Mendes on 9to5Mac:

In the study titled MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer, a team of nearly 30 Apple researchers details a novel unified approach that enables both image understanding and text-to-image generation in a single multimodal model.

Mendes gives a detailed explanation of the paper’s results – Apple’s image generation is “comparable” to GPT-4o in some tests, for example:

As a result of this approach, “Manzano handles counterintuitive, physics-defying prompts (e.g., ‘The bird is flying below the elephant’) comparably to GPT-4o and Nano Banana,” the researchers say.

We can basically see here that Apple’s models are a year behind where they want to be – but potentially catching up thanks to new research.

View the original.

Posts You Might Like

Apple’s Smart Home Display Now Coming This Fall: Here’s How to Get Ready Now »
Mark Gurman has reported Apple's smart home display won't ship until the fall to debut alongside the new Siri features – here's the developer session you can watch now to get prepared.
Apple Intelligence will support German, Italian, Korean, Portuguese, and Vietnamese in 2025 »
Apple Intelligence will now be rolling out to 16 more countries than the U.S. – with more coming too.
Exploring Conferences: Preparing for WWDC 2024! »
Rudrank Riyam has a great post summarizing his experience and goals for WWDC – reading this is getting me excited!
OverPicture gives easy access to Picture-in-Picture on macOS, iOS, and iPadOS »
OverPicture adds easy access to picture-in-picture from your Safari toolbar – get it for a few bucks on the App Store.