Skip to content
Metadata & enrichment

Complete book metadata — extracted, enriched, and verified

Most tools can only reshape the metadata you already have. Origami reads the actual EPUB or PDF, fills every gap with AI, and checks the result against the book’s own text — one clean, ONIX-ready record with every field sourced.

Titles Contributors Identifiers BISAC Thema Keywords Descriptions Language Accessibility Prices
Trusted by publishers worldwide
1,000+ organizations
50+ countries
Powered by the publica.la platform
Built on the standards: epubcheck DAISY Ace ONIX 3.0 BISAC / Thema EDItEUR List 196 EAA
The difference

Start from the book, not a feed

Most metadata automation starts from the ONIX you already have and makes it look better. It can’t tell you the declared language is wrong, that an ISBN doesn’t belong, or that an obvious subject is missing — because it never opens the book.

Origami does. It reads the EPUB or PDF itself, so enrichment is grounded in what’s actually inside, and every value carries its source: taken from the file, inferred by AI, or verified against the text. That’s the difference between a tidier spreadsheet and a record you can trust.

One record, every value sourced
ISBN-13
978-84-1234-567-8 From the file
Language
Spanish Verified
Contributors
2 · with roles AI
BISAC
FIC019000 AI
Thema
FBA AI
Keywords
24 terms AI
Accessibility
WCAG 2.1 AA Verified
Description
short + long Your edit
From the file AI Verified Your edit
Holistic by design

One pass, the whole record

Extraction, enrichment, verification and standards in a single run — not four tools stitched together.

  1. 01

    Extract

    Read every field already embedded in the EPUB or PDF — no re-keying.

  2. 02

    Enrich

    Fill the gaps with AI: subjects, keywords, contributor roles and descriptions.

  3. 03

    Verify

    Cross-check the record against the book’s own text and flag what doesn’t match.

  4. 04

    Standardize

    Normalize everything to ONIX for Books 3 with BISAC and Thema subjects.

  5. 05

    Publish

    Export the record, feed retailers and libraries, or write it back into the file.

Everything a record needs

Not a handful of fields — the complete bibliographic record the trade expects, in one place.

Identifiers

ISBN-13 and ISBN-10, product form and edition — correctly typed to the ONIX identifier lists.

Contributors & roles

Authors, editors, translators and illustrators, each mapped to an ONIX List 17 role code.

Titles & imprint

Title, subtitle, publisher and imprint — cleaned and consistently cased.

BISAC & Thema

Subject codes in both the North-American (BISAC) and international (Thema) schemes.

Keywords

Retail keywords in the language of the book, generated for discovery — not filler.

Descriptions

Short and long descriptions that follow the ONIX text-type conventions retailers expect.

Language

The declared language, checked against the prose and corrected when they disagree.

Accessibility

ONIX accessibility features (List 196) derived from the file — the metadata the EAA needs.

Prices & availability

For publica.la titles, live price and availability folded straight into the record.

Provenance

Every field is sourced — and your data always wins

Enrichment you can audit. Each value is tagged with where it came from, and nothing you decide is ever silently overwritten.

From the file

Read from the file

Whatever the EPUB or PDF already declares is taken verbatim and marked as embedded.

AI

Enriched by AI

Missing fields — subjects, keywords, roles — are inferred in the language of the book and clearly labelled.

Verified

Verified against the text

The declared language, ISBN and key subjects are checked against what the book actually contains.

Your edit

Edited by you, and kept

Any value you change is locked. Re-running enrichment never overwrites your edits.

Feed-only automation vs Origami

Both hand you cleaner ONIX. Only one of them actually reads the book.

Capability Feed-only tools Origami
Reads the actual book file (EPUB / PDF)
Fills genuinely missing fields with AI Partial
Verifies metadata against the text
Shows the source of every field
Your embedded data always wins
ONIX for Books 3 output
BISAC & Thema subjects
Accessibility metadata (EAA)
Writes corrections back into the file
Standards & compliance

Built on the standards the trade runs on

The record leaves Origami ready for retailers, aggregators and libraries — no re-mapping.

ONIX for Books 3

Every record is normalized to ONIX 3.0 (3.1-ready), the global standard maintained by EDItEUR.

BISAC & Thema

Subjects in both schemes, so a title is shelved correctly in every market.

EDItEUR code lists

Roles, languages, countries and product forms drawn from the official code lists (Issue 73).

Accessibility (EAA)

ONIX List 196 accessibility metadata — required to sell ebooks in the EU under the European Accessibility Act.

How it works

1

Upload a file

Drop in an EPUB or PDF from your workspace.

2

We read, enrich and verify

Origami reads what’s embedded, fills the gaps with AI, and checks it against the text.

3

Export or write back

Review the sourced record, feed it to your channels, or write it into the file.

Questions about metadata enrichment

How is this different from a metadata feed tool?
Feed tools reshape metadata you already have. Origami reads the book itself, so it can fill genuinely missing fields and verify the rest against the text — not just re-format a spreadsheet.
How is this different from Origami’s metadata extraction?
Extraction pulls what’s already in the file. Enrichment goes further — it fills the missing fields with AI, verifies them against the text, and completes the ONIX record.
What formats can it read?
EPUB and PDF.
Does AI overwrite the metadata in my file?
No. Anything embedded in the file always wins, and any value you edit is locked — AI only fills what is genuinely missing.
What does “verify against the text” mean?
Origami compares the declared metadata with the book’s own content — flagging, for example, a language tag that doesn’t match the prose or an identifier that doesn’t belong — so errors surface before they reach retailers.
Will it produce accessibility metadata for the EAA?
Yes. Origami derives ONIX List 196 accessibility features from the file — the metadata the European Accessibility Act requires to sell ebooks in the EU.
Is my content used to train AI?
No. Origami’s AI providers do not use your content to train models — see our AI policy for the details.

Metadata is normalized to the ONIX for Books 3 standard maintained by EDItEUR, using code lists Issue 73. BISAC is a trademark of BISG; Thema is maintained by EDItEUR. Accessibility features follow ONIX List 196. AI-inferred fields are drafts — review before publishing.

Enrich your catalogue’s metadata in Origami

Start free and build a complete, verified record for your first title.