Shrinking Chromium's Binary, One Table at a Time

โšก Chromium ๐Ÿ”ง C++ / Binary Size ๐Ÿ‘ค Helmut Januschka

Sometimes the expensive part of generated code is not the data. It is spelling out the same operation once for every entry.

Status: ๐Ÿšง Ongoing

The Shape of the Problem

I previously wrote about moving static data out of the Chrome binary - putting USB IDs and compressed phone metadata into lazily-loaded resources. This is the other side of binary-size work: data that legitimately belongs in the binary, but is represented by far too much generated code.

Chromium generates a surprising amount of C++: feature registries, UKM maps, enum-name lookups, sanitizer configurations, color mappings, locale caches, and resource maps. A generator often emits the most direct implementation - one initializer, function call, or setter per entry.

That is readable in the generator, but expensive in the binary. Repeating an operation for every entry produces machine code, relocations, padding, and sometimes static constructors. When the operation is identical and only the inputs change, the implementation wants to be data plus one loop.

Measuring the Whole Batch

I built an interactive binary-size dashboard from Chromium's android-binary-size and compile-size trybot results. Every point links back to its CL and records the newest patchset for which a size bot reported numbers.

The dashboard can switch between per-CL deltas and a cumulative view, include or exclude unlanded CLs, and apply a deliberately conservative rule: count savings from an unlanded CL, but do not pretend its size increases have shipped.

For the 12 table/offset changes discussed here, the latest recorded patchsets add up to roughly:

Those are not interchangeable measurements. Architecture, alignment, relocation encoding, and linker decisions can make the same representation win by very different amounts. The dashboard keeps the raw dimensions separate instead of collapsing them into one impressive-looking number.

Case Study: GL Enum Names

The generated GL enum table stored a std::string_view in every entry. On 64-bit builds that made each entry 24 bytes and required a relative relocation for every name.

The replacement concatenates all names into one string blob and stores a 32-bit offset:

struct EnumName {
  uint32_t value;
  uint32_t name_offset;
};

The entries shrink from 24 bytes to 8, the table moves from .data.rel.ro to .rodata, and the per-entry name relocations disappear. The current size-bot patchset reports 18,904 bytes saved on arm64, while arm32 grows by 1,684 bytes. That split is useful evidence: a source-level representation change does not have one universal binary-size result.

Case Study: Workaround Flags

GpuDriverBugWorkarounds::ToIntSet() expanded an if plus push_back for every workaround. The result was many copies of the same instruction sequence.

A constexpr table of workaround IDs and pointers-to-members lets one loop do the work. The output vector is unchanged; the current dashboard result is 2,160 bytes saved on arm64 and 1,760 bytes on arm32.

Case Study: Generated Registries

The same pattern appeared in several generators:

Representative landed changes:

The same experiment is still in review for Trusted Types event-handler names, settings dispatch, color-mixer registrations, sanitizer builtins, and default share-ranking data. Together they test where table-driven startup remains clearer and where the extra indirection is not worth it.

Resource Maps: Keep the Development Feature

GRIT resource maps had a different trap. Each production entry carried a path pointer, ID, and an std::optional filesystem path used only by a local WebUI development workflow. Removing the development path would save space but break a useful workflow.

The compact representation keeps both behaviors: URL and optional development paths become offsets into one deduplicated string blob. Production entries shrink without deleting the local loading feature. The current trybot result reports 31,072 bytes saved on arm64 and 6,404 bytes on arm32.

Finding Candidates

The useful review question is:

Does this generator emit one copy of an operation for every entry, when it could emit the entries as data and execute the operation once?

Good signs are generated switches, long initializer lists with non-trivial constructors, one static variable per feature, repeated string pointers, and large tables in relocation-heavy sections.

Binary size is rarely one spectacular deletion. It is a collection of representation choices. Tables, offsets, and loops are not automatically better, but when behavior is uniform and inputs are static, they usually express the problem more directly - and the linker tends to agree.