Depth-Aware Hybrid On-Device/Cloud Perception for Controllable Mobile Photo Editing
Abstract
Photo editing on a phone is moving from sliders to edits that understand the scene. Re-focusing on a subject, or inserting an object behind existing geometry, needs a depth map and a subject mask for every image. The generative models that carry out such edits need far more memory than a phone has. Where each part runs is therefore the central design question, and sending everything to the cloud makes every tap slow and puts every photograph on a server. We present DEPTHEDIT, a mobile editor built on a hybrid split. Segmentation and depth run on the phone’s NPU. Text grounding and diffusion run in the cloud and receive only the small crop they modify. The phone keeps one set of perception maps that every editing feature reuses, and refreshes them only when an edit changes the scene. We then measure what the split buys. Caching the segmentation model’s image embedding cuts a repeat tap from 277 ms to 4 ms, and the masks do not change. A ∼600 MB depth model replaces a ∼12 GB one with no measurable loss in the depth-of-field renders it produces. Against three alternative deployments, the hybrid uploads 86% less than a cloud editor, taps fastest, and never leaves a full photograph on a server beyond one request. The refresh rule earns its place too. A stale depth map mis-focuses 31% of a removed object’s pixels, and correcting it costs one 47 ms pass.