Identify / specification
Title Screen identification: the canonical form (canon1) and the set hash (set1)
Status: draft 2026-10-07, version 1 (revised 2026-10-07 after the first scans of real copies; the revisions are listed in section 11). This document is part of the Title Screen contract. Anyone can implement it in any language from this text and the public dump lists alone; the reference implementation (gamedb.ident, python -m gamedb ident) exists to be checked against, not to be required. Once the first public archive ships, this version is frozen: any change to a rule below is a new form (canon2) with its own column, never a silent edit.
1. Purpose
Given any copy of a game, however it is named, wherever it sits, and whatever container it is packed in, produce a value that names which version it is: platform, region, revision. The value is looked up in the Title Screen archive, which returns the dump, the release or releases it was sold as, and the work.
What this is not: a tamper check. A copy with one changed sprite gets a different value and is simply not identified. Forgery is out of scope; the form is chosen for compatibility and speed, not security.
2. Terms
- File: a sequence of bytes after every container has been opened.
- Dump: an entry of a dump list (No-Intro, Redump, MAME, FBNeo, the MAME software lists, TOSEC): one or more files that together are one version of one game. Title Screen's
dumptable has one row per entry. - Canonical file: a file after the transformation in section 4 or 5.
- Identity: the pair (
sha1,size) of the canonical file, or of the set (section 6) when the dump has several files.sizeis in bytes; for a set it is the sum of the files' sizes. Hex is written lowercase and compared case-insensitively. - Pre-check:
crc32+sizeof the canonical file. A ZIP archive stores each member's CRC32, so a member can be checked without inflating it. A pre-check hit is a candidate, confirmed by the identity. - Companion hashes:
md5of the canonical file (section 8) and, reserved,sha256. Neither is an identity.
3. Step one: open containers
Containers are opened recursively until no member is itself a container. Recognised containers: ZIP, 7-Zip, RAR, gzip, bzip2, xz, zstd, tar. Rules:
- A member is a file. Its name and path are used for one thing only: pairing a cue sheet or GDI with the track files it names (section 5). Names never affect the hash.
- A container that holds exactly the files of one dump identifies as that dump. A container holding several games yields several identifications.
- The container's own bytes are never hashed for identity. The only use of container metadata is the ZIP member CRC32 as a pre-check.
- Encrypted or unreadable members are skipped and reported, not guessed.
4. Step two, cartridges and digital files: strip the dumping tool's header
The canonical file is the ROM as the hardware sees it, which is how No-Intro and the MAME software lists hash it. Only the rules below transform bytes; every other platform's file is canonical as is.
| platform | detection | transformation |
|---|---|---|
| NES, Famicom | first 4 bytes 4E 45 53 1A (NES\x1a, iNES or NES 2.0) | drop the first 16 bytes |
| Famicom Disk System | first 4 bytes 46 44 53 1A (FDS\x1a) | drop the first 16 bytes |
| SNES, Super Famicom, Satellaview | size mod 1024 = 512 | drop the first 512 bytes (copier header) |
| Nintendo 64 | first 4 bytes 37 80 40 12 | swap each pair of bytes (.v64 byte-swapped form to big-endian) |
| Nintendo 64 | first 4 bytes 40 12 37 80 | reverse each 4-byte word (.n64 little-endian form to big-endian) |
| Nintendo 64 | first 4 bytes 80 37 12 40 | none (.z64 big-endian is canonical) |
| Atari Lynx | first 5 bytes 4C 59 4E 58 00 (LYNX\0) | drop the first 64 bytes |
| Atari 7800 | first 10 bytes 01 41 54 41 52 49 37 38 30 30 (\x01ATARI7800) | drop the first 128 bytes |
| PC Engine, TurboGrafx-16, SuperGrafx HuCard | size mod 131072 = 512 | drop the first 512 bytes |
| Master System, Game Gear, SG-1000 | size mod 16384 = 512 | drop the first 512 bytes |
| Genesis, Mega Drive, 32X, Pico | size mod 16384 = 512 and bytes 8 and 9 are AA BB (Super Magic Drive .smd) | drop the first 512 bytes, then de-interleave each 16384-byte block: the block's first 8192 bytes are the odd output bytes, its last 8192 the even output bytes (out[2i+1] = block[i], out[2i] = block[8192+i]) |
Notes:
- The rules are byte rules, not extension rules. An
.sfcwith a copier header is stripped; a.smcwithout one is left alone. Extensions only suggest the platform when the bytes are ambiguous. - UNIF (
UNIFmagic) NES files and other dumping-tool formats not listed are not canonical; they are reported as unidentified by form. - Game Boy, Game Boy Color and Advance, DS, 3DS, Virtual Boy, WonderSwan, Neo Geo Pocket, Atari 2600 and 5200, Jaguar, ColecoVision, Intellivision, Vectrex, Odyssey 2, Channel F, arcade ROM files, Switch
.xciand.nsp, PlayStation.pkgand.pbp, Wii.wad: the file as is. - Digital packages are canonical as is, and their headers name what they are; an implementation reports these as the header identity beside the hashes (informative, like FP1.md section 5): a PS3 or PSP
.pkg(7F 50 4B 47): its type at bytes 6 to 7 (1 PS3, 2 PSP or Vita) and its content id at bytes 48 to 83, whose characters 7 to 15 are the title id; a PS4.pkg(7F 43 4E 54): the content id at bytes 64 to 99; a.pbp(00 50 42 50): the PARAM.SFO it points to at byte 8 (CATEGORYMEis a PlayStation title packaged for PSP,DISC_ID,DISC_VERSION). - Some lists hash a form other than canon1 for some platforms (a headered set, an encrypted DS set beside a decrypted one, a trimmed Switch card). The archive records the list's own hash as a second row of
file_hashwithform = file, so a lookup by the whole-file hash also answers. An implementation computes both the whole-file hash and the canonical hash in one pass; they differ only when a rule above fired.
5. Step two, discs: the dump list's image form
5.1 CD-ROM family: Redump track files
Platforms: PlayStation, Saturn, Sega CD and Mega CD, PC Engine CD and TurboGrafx-CD, 3DO, CD-i, Neo Geo CD, PC-FX, Amiga CD32, Jaguar CD, Naomi and Triforce discs in CD form, and any other CD-based system.
The canonical form is Redump's: one file per track, 2352 bytes per sector, data tracks raw (sync, header, user data, EDC and ECC as on the disc), audio tracks as 2352-byte PCM frames, subchannel data excluded. Track boundaries and order come from the cue sheet.
- A dump is the ordered list of its track files. Its identity is the set hash (section 6) over the tracks' SHA-1s in cue order, and its size is the sum. A disc with a single track is a single-file dump: its identity is that file's own SHA-1.
- The cue sheet is not one of the files. Redump's lists name the
.cuebeside the tracks; it is metadata (names and track layout) and never enters the identity, so a copy whose cue was rewritten by another tool still identifies. - Sources that reach this form losslessly:
.cue+.binfiles as held; a CHD made from them (the tracks are reconstructed with chdman'sextractcd; where chdman writes one.binfor the disc, it is split at the cue'sINDEXsector positions into per-track files); a.binwith 2448-byte sectors (drop the last 96 bytes of every sector). - Sources that do not reach this form and are identified by the fingerprint only (
docs/spec/FP1.md): a 2048-byte-per-sector.isoof a CD title (EDC, ECC and subheaders are gone and cannot be assumed), a CHD whose data tracks are stored as 2048-byte sectors (made from such an.iso, or from a GDI with 2048-byte tracks), CloneCD.img, Alcohol.mdf, Nero.nrg, DiscJuggler.cdi. A later version may add lossless splitting rules for some of these. - A raw image of a whole multi-track disc with no cue sheet (one
.binholding every track) cannot be split without its sheet: track lengths are not recorded in the data. It is not identifiable by this form or by the fingerprint until its sheet is supplied. A raw image of a single-track disc is that track. - Known caveat: CHDs made by some chdman versions place a track's pregap differently from Redump's cue, so the reconstructed tracks differ in size from the list's. The archive marks dumps for which the round trip was verified; the fingerprint still identifies them.
5.2 GD-ROM: Redump track files
Platforms: Dreamcast, Naomi GD-ROM, Triforce GD-ROM.
The canonical form is Redump's track files: one file per track, 2352 bytes per sector, subchannel excluded, in track order. Redump lists them with a cue sheet whose REM HIGH-DENSITY AREA line opens track 3; the older .gdi sheet names the same files in the same order. Either sheet is metadata and never enters the identity. Identity is the set hash over the tracks in track order.
- A
.gdiset with 2048-byte data tracks (common in TOSEC-style sets) does not reach the form: fingerprint only. A.cdiis fingerprint only. - A CHD of a GD-ROM is reconstructed with chdman to per-track files. Its round trip to Redump's file boundaries is verified per dump: the first scans found CHDs whose last data track came back 225 sectors shorter than Redump's (the pregap that precedes a data track after audio), which the archive marks as
round_trip = mismatch; the fingerprint still identifies them.
5.3 DVD, UMD, GameCube, Wii, Xbox, Xbox 360, Blu-ray: the full image
Platforms: PlayStation 2 (DVD titles; PS2 CD titles follow 5.1), PSP, GameCube, Wii, Xbox, Xbox 360, PlayStation 3, Wii U.
The canonical form is the full disc image in 2048-byte sectors as the list hashes it (.iso, .wud for Wii U). A dump is one file and its identity is that file's SHA-1.
- Lossless containers that decode to the exact image: RVZ, WIA and GCZ (Dolphin's formats; Dolphin's own conversion reproduces the image byte for byte), CSO version 1 and 2 and ZSO (PSP; zlib or LZ4 blocks with an index table, the uncompressed-block flag honoured), WUX (Wii U), a CHD made with chdman's
createdvd, a CHD made with chdman'screatecdfrom the image (oneMODE1/2048track whose bytes are the image), and any ZIP or 7-Zip of the image. - Not canonical, fingerprint only: a scrubbed GameCube or Wii image (unused areas zeroed), an Xbox or Xbox 360 image reduced to its game partition (XISO), a PlayStation 3 image that was decrypted (Redump hashes the encrypted image). These are common in the wild, so the fingerprint specification handles each of them explicitly.
- Multi-disc games: each disc is its own dump, as the lists have it. A release with several discs has several dumps.
5.4 Arcade sets
A MAME or FBNeo set is a dump of several files; the canonical file is each ROM image as is. The set's identity is the set hash over the files sorted by their list names (bytewise, ascending), because an arcade set has no natural order. CHDs inside sets are hashed by the CHD's own internal SHA-1 as MAME lists it.
6. The set hash set1
For a dump of several files in canonical order (track order for discs, the cue sheet itself excluded; name order for arcade sets; list order otherwise):
set1 = SHA-1( hex(sha1_file_1) + hex(sha1_file_2) + ... )
where each hex() is the 40-character lowercase hexadecimal SHA-1 of a canonical file and + is ASCII concatenation with no separator. Order matters. A single-file dump has no set hash: its identity is the file's SHA-1, so a plain sha1sum matches it.
Test vector:
sha1("a") = 86f7e437faa5a7fce15d1ddcb9eaeaea377667b8
sha1("b") = e9d71f5ee7c92d6dc9e92ffdad17b8bd49418f98
set1([a, b]) = 5463504435e4dbf2b93a3a8a00ca78e36ea40e24
set1([b, a]) = b7d99985b3cf2b2e59215451e8b633a6671bd533
7. When two releases are one dump
When a publisher shipped identical bytes in two markets, no hash can tell them apart, and Title Screen does not pretend to. No-Intro lists one dump, "Super Metroid (Japan, USA) (En,Ja)", for the Japanese and the US cartridge; the archive holds one dump row that belongs to both releases (dump_release). A lookup returns every release the dump was sold as, and the API marks the answer identical_releases: true. Consumers that need one market must decide from something other than the bytes. A dump the list names "(World)" ("Super Mario Bros. (World)") is one release in the market WORLD: the list does not say which markets it reached, and the archive does not guess.
8. Companion hashes
crc32: the pre-check (section 2). Four billion values are not enough to be an identity across a few hundred thousand dumps.md5of the canonical file: shipped for cartridge platforms because it equals the RetroAchievements game hash there (their rules strip the same headers and convert N64 to big-endian). For disc platforms RetroAchievements hashes selected parts of the filesystem instead; that is a different scheme and, if ever added, a separate form (ra1).sha256: reserved column, not computed in version 1.
9. Test vectors for canon1
iNES file: "NES\x1a" + 12 zero bytes + (bytes 0x00..0xFF) repeated 64 times
whole file size 16400 sha1 91b79a50d2afa6f5c53cb9692a25cb9ef56a7007
canon1 size 16384 sha1 80cb9c430d80c3084649f65e0ca25dabbffb1b62 crc32 e81722f0
N64: the 64-byte file 80 37 12 40 04 05 06 ... 3F
z64 sha1 8be2f65de9da330a6e9b697beec04616b9fb4c36
the same ROM as v64 begins 37 80 40 12, as n64 begins 40 12 37 80; both canonicalise to the z64 sha1
SMD: a 32768-byte raw ROM whose byte i is (i * 7) mod 256
raw sha1 ee33e33e8e3a66205ad7ee4570116ed19f66301e
as .smd size 33280: a 512-byte header, all zero except bytes 8 and 9 = AA BB, then the two
16384-byte blocks interleaved (each block: the raw block's odd bytes, then its even bytes)
whole-file sha1 f17eec13dbe6a9d90749dff2dd07eb8673492a15; canon1 sha1 = the raw sha1
The script that produced these vectors is 40 lines of Python using only hashlib, struct and zlib; the reference implementation's unit tests regenerate them from the same synthetic inputs.
10. What a lookup returns
Input: any of sha1 + size, crc32 + size, fp1, or a file. Output, per candidate: the dump (list, name, platform, region, revision), every release it belongs to, the work, and the basis of the match: sha1 (identity confirmed), crc32 (pre-check only, not confirmed), fp1 (fingerprint, see its spec), file (matched a list's whole-file hash rather than the canonical form). Nothing about the caller's file is echoed back except the values it asked about.
11. Revisions to the draft
Version 1 is a draft until the first public archive. These clarifications came from scanning real copies (2026-10-07) and change no hash a conforming implementation already produced, except the corrected test vector.
- Section 9: the
.smdwhole-file vector was wrong (dc42a198...cannot be produced from the construction described); the header is now spelled out and the value corrected. The canon1 value was right. - Section 5.1: the cue sheet is not one of a dump's files; a CHD with 2048-byte data tracks is fingerprint only; a raw multi-track image without its sheet is not identifiable.
- Section 5.2: Redump lists GD-ROMs with a cue sheet; the
.gdinames the same files. The CHD round-trip mismatch is described as measured. - Section 5.3: a
createcdCHD of a DVD image is lossless. - Section 4: the header identity of digital packages (
.pkg,.pbp). - Section 7: the example is a "(Japan, USA)" dump; a "(World)" dump is one release in the market
WORLD, as the archive models it.