Swin2SR · ONNX Runtime Web

AI Image Upscaler

Enlarge a low-resolution photo with a super-resolution neural network that runs in your own browser. Your image is never uploaded — only the model comes to you.

Drop an image to upscale

Drag a file in, paste from the clipboard (Ctrl+V), or pick one. JPEG, PNG, WebP and GIF are supported.

Drop your image here
or click to browse · nothing is uploaded, ever
Runs on WebGPU where available WebAssembly fallback No account needed
Settings
Your pixels stay on this device. There is no upload request in this page’s code. The network tab shows only one-time downloads of the open-source model and library — your image is never in any of them.
Scale
Output 0 × 0
Fidelity 100% AI
Full neural detail. Drop toward 0% to blend back to the pixel-exact original.
Sharpen 35%
Edge contrast on the final image. 0% leaves it untouched.
Model precision
Picks the weight file that suits your hardware.
Preparing…
Result

Drag the handle to compare. The left side is your original, the right side is the output.

Source thumbnail
Upscaled result
Original before upscaling
Original Upscaled

How browser-based image upscaling works

An image upscaler has one job: given a grid of pixels, produce a bigger grid that still looks like the same picture. The naive way to do that is to interpolate — estimate each new pixel from the neighbours around it. It is cheap, instant, and it never invents anything, which also means it never adds the detail you were hoping for.

The tool on this page offers both approaches, and lets you decide how much of each you want.

AI Enhance runs a super-resolution network called Swin2SR. It is a vision transformer trained specifically for this task: given a small image, predict what a larger version of that image should contain. The network was converted to ONNX, a portable format, and is executed inside your browser by ONNX Runtime Web. Where your machine exposes WebGPU the computation runs on the graphics processor; otherwise it falls back to WebAssembly on the CPU. In both cases the model file is downloaded once and then runs entirely on your own hardware.

Accurate Resize uses no model at all. It resamples with the browser’s high-quality built-in resampler and applies an optional sharpening pass. Nothing is estimated, so no structure can be altered — which is exactly what you want for text, screenshots and diagrams.

What “fidelity” is for

A super-resolution model cannot recover detail that was never captured. Once pixels are gone they are gone, and what the network produces is an estimate built from patterns it learned during training. On a photograph, that estimate is usually convincing. On a document, a logo, a face, or anything where an exact value matters, the estimate can be quietly wrong — it may reshape a glyph, thicken a line, or imply texture that was never there.

The fidelity slider makes that trade explicit. At 100% AI you get the network’s full output. Drag it toward 0% and the result blends back toward plain interpolation, which changes no detail it was not given. Somewhere between the two is usually the right answer for the image in front of you, and the before/after slider above is there so you can judge it rather than trust it.

Choosing an engine
Image typeSuggested engineWhy
PhotographsAI EnhanceEstimating texture is desirable; noise and grain benefit from it.
Grainy or compressed shotsAI EnhanceThe network is trained to clean up exactly this kind of degradation.
Screenshots and UIAccurate ResizeInterpolation keeps every edge where it belongs.
Documents and scanned textAccurate ResizeA network may reshape letterforms, which changes meaning.
Logos, line art, diagramsAccurate ResizeCrisp geometry is reproduced exactly, not approximated.
Barcodes and QR codesAccurate ResizeModule boundaries must stay pixel-sharp to remain scannable.
Mixed photo and textAI Enhance, low fidelitySome neural detail without disturbing the type.

Measured against plain resampling

To check that the model is genuinely doing more than interpolation, the same photograph was taken at a known high resolution, degraded to a quarter of its size to simulate a low-resolution upload, then restored twice — once by plain resampling, once by Swin2SR. Both outputs were then compared pixel-for-pixel against the original they were trying to reconstruct.

Mean absolute error against the original (lower is closer)
MethodErrorEdge detail
Plain resampling, 4×2.5911.84
Swin2SR, 4×0.9513.12
Original (reference)—11.69

The neural output landed roughly 2.7× closer to the true original than interpolation did, while carrying measurably more edge structure. Note that the reference itself sits below the model on edge detail: the network emphasises edges beyond the original rather than merely reproducing them, which is why the fidelity slider matters.

These numbers are one test on one photograph. The gain is real and repeatable in direction, but its size depends entirely on the image. A clean, well-lit photo benefits more than a noisy one; and on text the model can score worse than plain interpolation, because inventing plausible-looking detail in a glyph counts as an error.

Privacy: where your image actually goes

Your image is read from the file you choose into a canvas in this page, then into memory, and then processed. The application contains no upload endpoint, no form action, and no code path that transmits pixel data. The requests this page does make are ordinary downloads: the inference library, the model weights, and the site’s privacy-friendly page-view analytics.

Limits worth knowing

A browser tab has a finite memory budget, so output is capped at roughly 8 million pixels — about 2800×2800. That is a 4× upscale of a 700×700 source, or a 2× upscale of a 1400×1400 one. Larger images are handled automatically in tiles, each one given surrounding context from the original so the seams between tiles do not show.

A single very large upload will be refused with an explanation rather than crashing the tab: either pick 2× or crop first. The model is fixed at 4× natively, so the 2× option runs the network and then resamples down, exactly as the reference implementation of that model does.

Frequently asked questions

Does this upscaler upload my image to a server?

No. The pixels never leave your device. There is no upload endpoint in the application code. The only outbound requests are downloads of the inference library, the model weights and the site’s page-view analytics — none of which contains any part of your image. The network tab shows exactly this: downloads before you run it, and nothing during.

Is the upscaling really done by AI inside my browser?

Yes. Swin2SR loads as an ONNX model file and executes in the browser via ONNX Runtime Web. On a WebGPU-capable GPU the work runs on the graphics processor; otherwise it falls back to WebAssembly on the CPU. Both run locally, which is the difference between this and an online upscaling service: the model came to your browser instead of your image going to theirs.

Can AI upscaling recover detail that was never captured?

No, and no upscaler can. What the model does is estimate plausible detail from what remains, guided by patterns learned from training data. On photographs that estimate is usually close to the original. On text, logos or faces it can be subtly wrong, because it is inventing rather than recalling. The fidelity slider exists for exactly this reason: at 0% you get plain interpolation that alters nothing.

When should I use AI Enhance instead of Accurate Resize?

Use AI Enhance for photographs, textures and natural scenes where new detail is desirable. Use Accurate Resize for screenshots, documents, diagrams, logos, line art, barcodes and anything where a specific pixel value matters, because that mode resamples what already exists and adds nothing. A good default is AI Enhance at moderate fidelity, checked against the comparison slider.

How large an image can this handle?

Output is capped at about 8 million pixels, roughly 2800×2800, because a browser tab has a finite memory budget. That corresponds to a 4× upscale of around a 700×700 source, or a 2× upscale of around a 1400×1400 one. Images are processed in tiles with surrounding context, which keeps peak memory low and is invisible in the output.

Why is the first run slower than later ones?

The first run downloads the inference library and the model weights, a few tens of megabytes depending on the precision you choose. Those are fetched over plain HTTPS from public CDNs and cached by your browser under their standard HTTP cache headers, so later visits reuse them. The image itself is never part of any download.

Does it work on a phone?

Yes. Accurate Resize needs no model and works immediately anywhere. The AI mode needs WebAssembly, which is essentially every modern browser; where WebGPU is also available the model runs on the GPU and is several times faster. The output pixel cap applies everywhere, so a very large phone photo may need 2× rather than 4×.