Why Image Generation Needs More Than Bigger Models with Fatih Porikli
EPISODE 773
|
AUGUST
12,
2026
Watch
Follow
Share
About this Episode
Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution image generated locally, and today’s models still struggle in surprising ways.
In this episode, Fatih Porikli, Vice President of Technology at Qualcomm, joins me to discuss what remains unsolved in image generation and several approaches his team presented at CVPR to address those challenges. We explore why better training objectives can improve controllability, how separating scene planning from rendering may lead to more reliable image generation, techniques for generating 16-megapixel images efficiently on edge devices, and new methods for eliminating the visible artifacts that often appear in AI-powered image editing.
Along the way, we discuss reinforcement learning for image generation, agentic image generation pipelines, on-device AI, and what the next phase of progress in generative vision systems is likely to look like.
About the Guest
Fatih Porikli
Qualcomm
Resources
- Resolving the Identity Crisis in Text-to-Image Generation (CVPR 2026)
- Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation (CVPR 2026)
- PixelRush: Ultra-Fast, Training-Free High-Resolution Image Generation via One-Step Diffusion (CVPR 2026)
- InvertFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting (CVPR 2026)
- Multi-Scale Local Speculative Decoding for Image Generation (CVPR 2026)
- ReHyAtt: Recurrent Hybrid Attention for Video Diffusion Transformers (CVPR 2026)
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer (CVPR 2026)
- PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference (CVPR 2026)
- FLUX.1 (Black Forest Labs)
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (introduces GRPO)
- CVPR 2026 (Conference on Computer Vision and Pattern Recognition)
- Distilling Transformers and Diffusion Models for Robust Edge Use Cases with Fatih Porikli - #738
- Gen AI at the Edge: Qualcomm AI Research at CVPR 2024 with Fatih Porikli - #688
- Data Augmentation and Optimized Architectures for Computer Vision with Fatih Porikli - #635
- Optical Flow Estimation, Panoptic Segmentation, and Vision Transformers with Fatih Porikli - #579
- Why Vision Language Models Ignore What They See with Munawar Hayat - #758

