Back
Pakistan's First Gen AI App
A brainstorm for a background generator app turned into a full AI creative suite, now with 25M+ downloads and models trained on 2B+ photos.

ROLE
Product Designer
OUTCOME
27M+ Dowloads
2B+ Images processed
DURATION
8 months
The Context
I wasn't hired to build an AI image generator.
I was deep in PhotoTune, an advanced photo editor for iOS, obsessing over sliders, masking tools, and export flows. Then, stakeholders dropped a brief on the table: "Let's brainstorm a background generator app."
Simple enough. Users hated bad backgrounds. We could generate better ones.
But somewhere in that brainstorm, a question changed everything: If we can generate the background… why are we still stopping at the background? If the model can imagine a beach, a studio, a neon-soaked alley, it can imagine the whole image.
That question became Imagine Art.

Listening to Someone Else's Users (Research)
We were entering a crowded space. Remini and other AI photo apps had already captured millions of users, so instead of guessing, we studied them.
But we didn't just audit their features. I dug into what their users were actually saying, App Store reviews, Reddit threads, support complaints, and social comments.
The feature lists told us what these apps did. The reviews told us where they hurt.
Three patterns kept surfacing:
The Blank Prompt Paralysis: Users opened the app, saw an empty text field, and froze. They wanted to create something, but didn't have the vocabulary to describe it.
Unpredictable Output: People typed something, got something wildly different, and quit. There was no sense of control or direction.
Prompting Was a Skill, Not a Feature: The best results went to power users who knew to write "8k, cinematic lighting, hyperrealistic."
Everyone else got mediocre images and blamed themselves.
That last insight became my north star:
The average person doesn't have a prompting problem. They have a starting problem.

Designing for the Blank Canvas
(The First Flow)
Most AI apps dumped users into a blank prompt box with a "Generate" button (a test that most failed). Imagine Art’s first flow followed one rule: never leave users facing an empty screen.
Three entry doors:
Write it: Clean, frictionless prompt field for users who already know what they want.
Pick it: Tappable prompt suggestions that drop editable text into the field (not one-tap generators). This turned passive users into remixers and taught prompting.
Get inspired: community feed of generated images so people could start from a result they liked.
Styles (Visual chips with live previews: Realistic, Anime, 3D, Cinematic, etc.). Highest-leverage move: a user could type five plain words and still get a polished, intentional image because the style absorbed all the technical prompting. Complexity moved from the keyboard into the interface.

The Design Bets
A few decisions defined the product's DNA:
1. Suggestions that teach, not shortcut.
Pre-filling the input instead of instantly generating meant every user unknowingly went through a prompting tutorial on their first session. Retention starts with a first success.
2. Style as a visual decision, not a written one.
Nobody should have to know the word "bokeh" to get depth of field. Visual chips let users point at what they want.
3. Speed to first image.
Every added step before the first generated image was a chance to lose someone. I fought to keep the path from app open → first image as short as humanly possible, and pushed advanced controls to a secondary layer for users who wanted them.
4. Mobile-first, thumb-first.
This was born inside an iOS editor. Every primary action, prompt, style, generate had to live in the bottom third of the screen.
The Impact
27M+
Downloads
2B+
Photos used to train
the image generation models
5 secs
Time to value
Chapter 6: Learnings
What building an AI creative tool for beginners taught me.
The model was never the bottleneck the blank input field was.
Guidance is the product.
Design so users learn by doing.
Stay curious past the brief.
What I'd do differently:
Design the moment after generation sooner. We obsessed over the first image, but retention lives in the second one; refining, remixing, returning. Building that loop from day one would have compounded faster.