An AI editing interface can look simple: upload an image, type an instruction, and press Generate. The engineering contract behind that screen is not simple at all.
Different models can accept different numbers of source images. A resolution may be available for one task but not another. Some paths may start as a guest while others require authentication. Aspect ratios, credit costs, availability, and output options can all change independently.
If the front end treats those conditions as static feature flags, users eventually reach combinations that the current backend cannot honor. A more durable approach is to model the editor as a capability-driven workflow.
The first state is not “ready to generate.” It is “input accepted.”
For an image tool, that contract should answer a few questions before any model selection happens:
- Is the file type supported?
- Is the file within the current size limit?
- How many source images are present?
- Does this workflow expect one image or several?
- Can the browser preview the input and let the user remove or replace it?
In the example used here, the reviewed public interface accepts JPG, PNG, and WebP files up to 24 MB per file. Its Single Edit path expects exactly one source image, while multi-image behavior depends on the selected model and mode.
That distinction belongs in validation, not in an error message after the user has written a detailed prompt.
Treat capabilities as data
The next state is capability negotiation. Instead of hard-coding a collection of dropdowns, the interface should derive valid choices from the current capability response.
A useful capability record can include:
model
mode
minimum and maximum source-image count
supported aspect ratios
supported resolutions
authentication requirement
credit cost
availability
The exact schema will vary, but the principle is stable: a choice should only be offered when the current combination supports it.
This matters because UI controls are not independent. Selecting a different model may invalidate the chosen resolution. Moving from Single Edit to a fusion mode may change the source-image count. A 4K option may be valid for one model, ratio, or task and invalid for another.
When capabilities are treated as data, the interface can recompute the valid state after each meaningful selection. When they are treated as marketing copy, invalid combinations leak into the generation request.
Make invalidation visible
Dynamic forms need an explicit invalidation rule. Suppose a user selects a resolution and then changes to a model that does not support it. The interface has several options:
- silently replace the value;
- keep the invalid value and fail later;
- clear the value and explain why;
- select a new default and announce the change.
The first two are hard to trust. The last two are usually safer because they keep the visible state aligned with the request that will actually be sent.
This is a small UX detail with a large operational effect. It reduces avoidable failed jobs and gives support teams a clearer account of what the user selected.
Separate request acceptance from generation success
Pressing Generate should not collapse the rest of the workflow into a spinner.
A browser AI job benefits from distinct states such as:
input_ready
request_validated
queued_or_processing
result_ready
reviewed
downloaded
recoverable_error
The names are not important; the separation is.
Request validation confirms that the payload matches the current capability contract. Processing communicates that the job is underway without promising a precise completion time. Result ready means an output is available, not that it is correct. Review gives the user a place to inspect the result before download.
That final distinction is essential for generative editing. A technically successful response can still contain artifacts, misunderstand the instruction, or produce a result that does not fit the intended use.
Review is part of the product, not an afterthought
Traditional form flows often end when the server returns success. Generative tools need another state: human acceptance.
The result view should make it easy to compare the output with the source and the instruction. It should also preserve enough request context to support a sensible retry: selected mode, model, aspect ratio, resolution, and source-image set.
The interface should not promise perfect edits, exact prompt compliance, or artifact-free results. Those are review questions, not guaranteed system properties.
For sensitive or evidentiary images, review also includes a more serious question: did the edit change information that should remain visible? A clean-looking result is not automatically an honest or appropriate result.
Authentication is another capability
Login requirements should be modeled the same way as other capability constraints.
If an eligible flow can begin as a guest but a selected model requires an account, the user should learn that when choosing the model—not after uploading files and composing a prompt. The interface can preserve the in-progress request while presenting the authentication gate, but it should not pretend the gate does not exist.
In this example, some supported paths can start without signup, while some models and saved account features require Google authentication. That boundary can change, so it belongs in current capability data and product copy, not in a permanent promise that the entire tool is account-free.
Design the failure states before the happy path ships
Capability-driven interfaces still need good failures. At minimum, distinguish among:
- unsupported input;
- invalid source-image count;
- stale or unavailable capability;
- authentication required;
- insufficient credits or a changed cost;
- provider or generation failure;
- result unavailable for review or download.
Each state should tell the user what changed and what action is possible. “Something went wrong” discards information the application already has.
The interface should also revalidate before sending the request. Capability data may have changed since the form loaded. Client-side constraints improve the experience, but server-side validation remains authoritative.
A concrete browser workflow
The pattern above is visible in AI Photo Editor, which I work with. Its public workflow is upload, describe the edit, choose currently supported options, generate, review, and download.
The useful engineering lesson is not a particular model name. It is the decision to validate models, modes, aspect ratios, resolutions, source-image counts, credit costs, authentication requirements, and availability as a changing contract.
That contract supports several editing families—such as background changes, object erasing, eligible enhancement, image extension, restoration prompts, style changes, and model-dependent multi-image fusion—without implying that every option works with every other option.
The durable abstraction
An AI editor is easier to evolve when the front end does not ask, “Which features do we advertise?” but instead asks, “Which request is valid right now?”
The durable abstraction has five parts:
- validate the source input;
- negotiate current capabilities;
- build and revalidate a generation request;
- separate processing from human review;
- make download a deliberate handoff.
This structure does not eliminate model changes or provider failures. It gives those changes a defined place in the interface, which is what makes a multi-workflow AI product understandable rather than merely feature-rich.
How are you representing changing model capabilities in your own front end: a schema, a state machine, server-driven UI, or something else?