Why Image Generation Needs More Than Bigger Models with Fatih Porikli
August 12, 2026
AI Summary
5 min readIn 2024, you can ask almost any AI system for a picture and get something that looks remarkably good. But "looks good" and "is correct" are not the same thing. Ask for several different people, and the model may generate variations of the same face. Ask for a specific number of subjects, and it may ignore that detail. Push toward higher resolution, and speed and memory become constraints. Fatih Porikli, Vice President of Technology at Qualcomm, argues that the next frontier in image generation is closing the gap between plausible images and precise, controllable results. His team presented more than twenty papers at CVPR 2024, several of which tackle this problem by rethinking objectives and architectures rather than just scaling up.
Diversity as an Explicit Training Objective
A core observation driving Porikli's work is that existing text-to-image models have not really learned to create truly distinct identities. When asked to generate a group of people, the model often duplicates the same face. As Porikli puts it, the problem is not image quality—pixel-wise, the output looks realistic—but a missing objective. "Existing training objectives focus heavily on realism and matching the user prompt, but they don't explicitly encourage diversity between people."
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 Main Outline
- 2 (00:31) **The Core Challenge: Gap Between Plausible and Precise** - Images look good but often fail on specifics like identity, count, and composition.
- 3 (01:30) **Root Cause: A Single Model Solving Too Many Problems** - The complexity of composing a scene with multiple people, preserving identities, and rendering is overwhelming for a monolithic model.
- 4 (03:21) **The Three Unresolved Buckets: Control, Quality, and Efficiency** - Despite impressive progress, key research challenges remain in controllability, image quality, and on-device efficiency.
- 5 (05:56) **Paper 1: Disco - Diversity as an Optimization Objective** - The paper introduces a reinforcement learning approach to explicitly reward identity diversity in generated images.
- 6 (09:30) **Why Not Just Use Better Data? The Value of the Right Objective** - The discussion frames the problem as an optimization challenge, not a data scarcity issue, allowing for efficient fine-tuning.
- 7 (11:41) **Broader Lesson: The Impact of Curriculum Learning** - Curriculum learning is now a practical tool for making multi-modal models better.
+ Full timestamped outline available in the app
Show Notes
Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution image generated locally, and today’s models still struggle in surprising ways. In this episode, Fatih Porikli, Vice President of Technology at Qualcomm, joins me to discuss what remains unsolved in image generation and several approaches his team presented at CVPR to address those challenges. We explore why better training objectives can improve controllability, how separating scene planning from rendering may lead to more reliable image generation, techniques for generating 16-megapixel images efficiently on edge devices, and new methods for eliminating the visible artifacts that often appear in AI-powered image editing. Along the way, we discuss reinforcement learning for image generation, agentic image generation pipelines, on-device AI, and what the next phase of progress in generative vision systems is likely to look like.
🗒️ Full show notes: https://twimlai.com/go/773
More from this podcast
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) →