ASCII FX — structural ASCII rendering for the web

switching…
loading profile…
Matcher
Emoji palette

chromatic-v1 takes the lowest error over whatever glyphs it is given, and a flat swatch is the best possible reconstruction of a flat cell. So solid blocks win constantly. Removing a glyph is the only way to stop it being picked. 100 of 1301 start selected: the usage-curated set, within 1% of the whole pool at a thirteenth of the per-cell work.

Source
Grain

Not a renderer option: the noise is painted onto the source before the matcher ever sees it, so it changes which glyphs get picked rather than tinting the ones already chosen. Turning it up makes a still image live — every frame is a full re-match.

Font
Backend
Matching
View
Interaction

Export

What is this?

Most ASCII filters measure how bright a cell is and pick a character off a ramp like .:-=+*#%@. That gets you surprisingly far, but brightness is the only thing it looks at, so a diagonal edge and a flat gray patch of the same brightness end up as the same character. Whatever shape was in the cell is gone before any glyph is chosen.

ASCII FX compares each cell against the pixel shapes of the font's glyphs and keeps whichever one rebuilds that cell most closely, then fits a color to it. Two constraints shaped everything below:

The pipeline

Some of the work only has to happen once. A compiler measures the font ahead of time and writes a profile: an 8×8 mask, a coverage value, and an atlas tile per glyph. The rest happens every frame, in four stages. What follows is one frame of that, running the same code the playground above is running. Hover any stage to follow a single cell through it, or drag the columns slider and watch the grid, the samples and every decision recompute.

Stage 3 colors each cell by its class: gray cells are transparent, blue are flat, orange carry structure worth matching. The sections below take the stages one at a time.

Compile the font

Ask a browser to draw a glyph and the shape you get depends on the machine. System rasterizers (browser canvas, CoreText, FreeType) all anti-alias differently, and determinism dies right there. So the compiler brings its own: integer scanline rasterization into a 64 px cell, 4×4 subsamples per pixel, fixed Bézier flattening, nonzero winding. Same font bytes, same profile bytes, always.

Per glyph it stores:

Below: the live atlas of whatever font is selected up top. Hover a glyph to see the 8×8 mask the matcher compares against. Notice how thin glyphs like _ and - get sparse or empty masks; those are reached through coverage, not shape.

What can go in a charset? Any code point the font has outlines for. The built-in ascii-blocks set adds half-blocks and quadrants (▀ ▄ ▌ ▚ ▓). Braille patterns (⠏ ⣠ ⣿) are a classic trick: each glyph is a 2×4 dot grid, which hands the matcher 256 distinct shapes. The compiler refuses characters the font lacks, so a compiled profile never contains tofu. Runtime profiles cannot check that; a missing glyph quietly profiles the fallback box. Color emoji fail all three of this page's assumptions at once. Most emoji fonts are bitmap-only with no outlines to rasterize, emoji advance two cells wide so the monospace check fails, and fitting one ink color per cell would repaint them anyway. That is why Emoji mode is a different algorithm rather than a bigger charset.

None of this applies in Emoji mode, because there is no font to rasterize. Colour emoji ship as PNG assets, so the compiler decodes them and box-filters straight down to the 8×8 descriptor and a square-celled RGBA atlas.

The profile records SHA-256 hashes of the font, charset, and parameters. Frames reference profiles by fingerprint — a short content hash that names this exact profile — so a frame can never decode against the wrong font.

The grid

Before anything is matched, the source has to be carved into cells. Text cells are taller than they are wide, so the cell aspect enters the row count. For source W×H and cell cw×ch:

Round half up, used everywhere in the pipeline.

The pipeline explorer above prints this exact formula with live numbers as you drag its slider.

Emoji mode changes nothing here. An emoji cell is square rather than tall, so the same formula simply lands on a different row count.

Shrink the source

Now every cell is boiled down to the same tiny fixed size, whatever the source resolution was. Each one gets exactly 8×8 = 64 samples, a box filter over the pixels underneath — a plain average of every source pixel under the sample, with transparent pixels weighing less (that is the alpha weighting below).

GPU texture filtering is banned here: bilinear taps round differently per vendor. The compute shader does the same integer sums, one workgroup per cell, one lane per sample.

Emoji mode shares this step exactly. Both matchers work from the same 64 reduced samples, and the only thing that differs is what they compare them against.

Classify each cell

The decisions from here on work on each sample's luma: its brightness collapsed to one 0–255 number, a weighted mix of red, green and blue. Green carries the most weight because eyes are most sensitive to it — pure green reads far brighter than pure blue at the same intensity:

Integer only. In the code that floor is a >> 8.

There is no flat path in Emoji mode. Nothing would be gained by one, since there is no coverage ramp to fall back to and no fitted colour that could degenerate, so every non-transparent cell takes the same route.

The source mask

Comparing a cell against 95 glyphs would be expensive if you compared pixels, so the cell is reduced to one bit per sample first. The darkest and brightest samples become two color endpoints, and every sample joins whichever one it sits nearer to by squared RGB distance. Because that comparison uses color rather than luma, a red stroke on equally bright green still masks cleanly.

Paint an 8×8 cell below and the matcher runs on every stroke: endpoint classification, polarity, the Hamming shortlist over the whole charset, and the exact rerank from the next two sections. The preloaded stroke is a soft-edged diagonal. Its gray edge samples join whichever endpoint is nearer, which is why anti-aliased strokes still mask cleanly.

Polarity: glyph masks say 1 = ink, so which source side counts as ink follows from the palette. White on black means the bright side is ink, and the mask is inverted before comparing. This is derived, never a flag. A free invert option would let the shortlist and the rerank optimize different orientations.

Emoji mode never builds a mask at all, and has no polarity to derive. An emoji's colour is baked in, so there is nothing to fit, and the objective becomes squared error against the emoji's own 64 samples composited over the backdrop it will be drawn on.

Shortlist by shape

The Hamming distance between two masks is simply how many of their 64 bits disagree — computed as two XORs and two popcounts (count-the-1-bits, one instruction on either backend). Keep the 8 smallest, ties to the lower glyph id; that shortlist is what the exact scorer reranks next. Scoring a glyph exactly costs about 64 multiplies, so the split pays for itself: cheap filter over the whole charset, exact score over 8. Same winner, a fraction of the cost. In full mode the score is min(d, 64−d), because free colors can flip orientation.

Fit colors, pick the winner

Partition the cell's 64 samples by the candidate's own mask, then fit:

glyph mask ──▶  ink samples        fg = their mean      (full, foreground)
                the rest           bg = their mean      (full)  or fixed backdrop
mono: both colors fixed by the palette

The mean is the unique minimizer of squared error for a fixed partition, so every candidate is scored with the best colors it could possibly have:

Worst case is 64 · 3 · 255² = 12,484,800, so the comparison fits in exact u32 math. Means are rounded to u8 before scoring, so backends cannot drift. Lowest error wins, ties keep shortlist order. The winner's id and colors are the frame.

Draw it

Drawing is the easy part by now. Each cell samples its glyph's atlas tile and blends fg over bg by coverage. The atlas is mipmapped so strokes keep their weight at any scale. The CPU backend builds the same mip chain in Canvas2D.

Interactions (reveal, push, magnify, glyph rotation near the pointer; wave and original-mix across the whole frame) warp this composite. They never re-run matching. The GPU evaluates them per pixel, the CPU per cell, same formulas.

Emoji mode samples an RGBA atlas instead of tinting a coverage plane. The colour is already in the glyph, so there is nothing to mix it with but the backdrop.

Bit for bit on GPU and CPU

Going faster

Temporal reuse, live. Red cells are the only ones re-matched this frame; everything else is served from the previous result because its samples are byte-identical:

Cheaper matchers

Both are opt-in and never chosen for you, because both are measurably worse than exact matching. The repo benchmarks say by how much.

Emoji mode has no prefilter at all. Over a curated palette of about 100 glyphs an exhaustive search costs roughly what structural-v1's shortlist and rerank cost together, and every shortlist measured lost more quality than it saved in time.

Glossary

The terms this page leans on, in one place. Each is also explained where it first appears above.

Performance

Picking glyphs by shape turns out to cost less main-thread time than every brightness ramp measured here, which is not the result you would expect. Every figure below is read out of a harness that regenerates them, so none of it can drift from what actually ran.

Versus other libraries

No npm package publishes image→emoji rendering. Every emoji-mosaic project we could find is an application, not a library, so there is nothing to install and bench. The comparison rows are reference implementations of the technique all of them share instead: one mean colour per emoji, one per cell, nearest wins. They run on the same source at the same square 160×90 grid with the same curated palette.

approachpicks glyphs bygridp50 msp95 ms~fps
ascii-fx · WebGPUshape + baked colour160×903.03.5333
mean-colour reference · cube LUTmean colour160×9011.011.791
mean-colour reference · linear scanmean colour160×9013.313.775
ascii-fx · CPU fallbackshape + baked colour160×9063.365.716

Speed is only half of it though. Mean-colour matching is a colour quantiser, and cannot see sub-cell structure at all, so the two are not doing the same amount of work for their time. Measured 2026-08-27, and regenerated with pnpm bench:compare.

ascii-fx · WebGPUshape-awareexact structural + fitted color
357 fps2.8 ms/frame p50
ascii-fx · no WebGPUshape-awareexact structural · workers + WebGL2
357 fps2.8 ms/frame p50
textmode.js 0.17brightness + color · WebGL
222 fps4.5 ms/frame p50
three.js AsciiEffectbrightness · DOM
122 fps8.2 ms/frame p50
aalib.js 2.0 · monobrightness · canvas
97 fps10.3 ms/frame p50
aalib.js 2.0 · coloredbrightness + color · canvas
60 fps16.8 ms/frame p50
chafa-wasm 0.3shape-awareshape-aware blocks · wasm
20 fps50.1 ms/frame p50

Real published packages, plus the two techniques this project credits as influences, all fed the same animated source at the same glyph grid. Over floor is what each approach costs you on top of simply drawing that source.

approachpicks glyphs bygridp50 msp95 msover floor
baseline · scene only, no ascii——2.73.0the floor
ascii-fx · WebGPUshape, exact160×422.83.2+0.1
ascii-fx · fallback, no GPU‡shape, exact160×422.84.7+0.1
ramp reference · mono*brightness160×423.03.4+0.3
textmode.js 0.17 (WebGL)brightness + color106×604.56.1+1.8
three.js AsciiEffect (0.185)brightness160×428.210.8+5.5
ramp reference · color*brightness + color160×429.410.2+6.7
aalib.js 2.0 · monobrightness160×4210.332.5+7.6
aalib.js 2.0 · coloredbrightness + color160×4216.828.6+14.1
ascii-fx · fallback, Canvas2D paint†shape, exact160×4231.935.0+29.2
ascii-fx ramp matcher†brightness160×4240.542.1+37.8
ascii-fx shape6-lut (Harri-style)†6-D shape vector160×4240.643.1+37.9
ascii-fx · fallback, one thread†shape, exact160×4242.044.3+39.3
chafa-wasm 0.3shape-aware blocks160×4250.158.7+47.4

Exact structural matching with a fitted colour per cell adds 0.1ms. Matching and compositing overlap on the GPU, so it lands below every brightness ramp in the table while doing considerably more.

Measured 2026-08-27 on one Apple-silicon laptop: headless Chromium, vsync off, 200 timed frames after 40 warmup, each library alone in a fresh page, best of two passes. * Not libraries. The brightness-ramp technique hand-written with zero overhead, marking the floor any ramp library could reach. † Main-thread JavaScript, composite-bound at this glyph size: matching alone is ~10.3ms against ~33ms for exact structural. ‡ No GPU at all: the exact matcher on worker threads, the grid painted by a WebGL2 fullscreen draw. Same cells as every other ascii-fx row — the three fallback rows differ only in where the matching runs and how the grid is painted.

Run it on your machine

The table above is a controlled measurement taken on one machine. Below is that same harness, with the same scene and grid, running in your tab instead.

One thing to watch for: your browser paints in step with your display, so anything that can hold your refresh rate will report much the same frame time. Read the result as pass or fail on your own hardware rather than as a ranking. The rows that fall below your refresh rate are the ones that cannot keep up, and the table above is where the ranking lives.

~30 seconds, 14 contenders

Loads about 8 MB of third-party libraries (three.js, textmode.js, aalib.js, chafa-wasm) on the first run. Nothing is fetched until you press the button.

What the fallback costs

The fallback runs the same matcher and produces the same output without a GPU. A still image is a one-off ~33ms at 160 columns; live video stays comfortable to about 120.

gridcellsp50 mscells/ms
80×211,68010.3163
120×323,84020.1191
160×426,72033.1203
240×6315,12070.5215
320×8426,880120.5223

What the cheap matchers buy

Both are opt-in, as above. This is how much they save, and what it costs you.

matcherspeed vs exactwhat you give up
shape6 + LUT3.3×small error deltas in full color; structural recall collapses on dense texture (worst corpus case 0.3%)
shape6 brute2.3×similar; no 512KiB LUT payload
ramp3.8×brightness only. An effect, not a fidelity mode

Using it

Compile a font once, hand the renderer a source, and everything above happens on its own. There are three ways in, depending on what you are building.

React

import { AsciiImage, AsciiVideo } from '@ascii-fx/react'

<AsciiImage src="/hero.jpg" alt="Portrait" columns={160} color="full"
  interaction={{ type: 'reveal', radius: 0.18 }} />

<AsciiVideo src="/clip.mp4" profile={profileRef} />

Anything else

const ascii = await createAsciiRenderer({ canvas, profile, backend: 'auto' })

ascii.setSource(imageOrVideoOrCanvas)
ascii.start()                          // rVFC for video, rAF otherwise
ascii.setOptions({ columns: 200, color: 'mono' })
ascii.setInteraction({ type: 'push' })

const frame = await ascii.captureFrame()

Three.js

const pass = new AsciiPass({ profile, renderer, columns: 160, color: 'full' })
await pass.init()

pass.render(scene, camera)             // per frame, instead of renderer.render()

// or in React Three Fiber
<AsciiEffect profile={profileRef} columns={180} interaction="reveal" />

Matching runs on Three's own device, so nothing is ever read back. WebGL is unsupported on purpose: exact matching there would need a GPU→CPU copy every frame.

Ahead-of-time profiles

export default defineAsciiConfig({
  profiles: { default: { font: './fonts/GeistMono.woff2', charset: 'ascii' } },
  frames:   { hero: { image: './src/hero.png', columns: 140, color: 'full' } },
})

import profileRef from 'virtual:ascii-profile/default'

Profiles compile once, cache by content, and hot-reload in dev, at about 13KB over the wire. Frames go further and precompute the matching itself, so hero art costs nothing at runtime. There is a CLI if you are not on Vite. This page eats its own cooking: the playground above loaded a build-time profile.

Narrowing the character set

// at build time
profiles: { rain: { font: './fonts/GeistMono.woff2', characters: '01 ' } }

// or narrow a built profile at runtime
const profile = subsetProfile(await loadProfile(profileRef), '01 ')

Subsetting carries each glyph's raster data over exactly, so a narrowed profile matches identically to one compiled that way. You do not have to recompile anything, and there is nothing for the two to drift apart on.