NER comparison

Compare word-list matching, GiNZA, XLM-RoBERTa, GLiNER2.5, Qwen3.5-4B, and GPT-6.1-Sol Med on the same text. The notebook interface and source text are Japanese.

Open the notebook

Open in your browser (University access only)

Run locally

Unpack the notebook bundle (ZIP, 104 KB) and keep all files together, including the hidden .marimo.toml. Run this command from that directory:

uvx --python 3.12 marimo==0.25.0 edit --sandbox ner-comparison.py

The notebook initially displays saved results. GPT-6.1-Sol Med is available only as saved comparison results. A new run downloads several GB of models on first use and processes the text on the CPU. No API key is required.

Compare results

  1. Inspect each method’s results for part 1 (four paragraphs) of Soranoha’s “The Spider’s Thread”.
  2. Switch to 「蜘蛛の糸:領域別カテゴリ」 (story-specific categories) and compare the extraction results for deities, sinners, and places of punishment.
  3. Select 「本文・条件を変えて実行」 (edit text and conditions), change one category name and description, and press 「比較する」 (compare).

You can inspect extraction criteria, provisional reference annotations (the previous GPT-6.1-Sol Med output), and raw outputs inside the notebook. If you edit the text or extraction criteria, the notebook displays extraction results without the reference-comparison table.

Word-list matching uses terms supplied for these particular texts. You can edit the terms and categories in 「本文・条件を変えて実行」.

GiNZA and XLM-RoBERTa use NER categories defined during training. The XLM-RoBERTa model was trained on annotated Japanese Wikipedia texts. This comparison uses its person (PER) and location (LOC) categories; other categories remain in the raw output.

GLiNER2.5 is also encoder-based, but accepts category names and descriptions at run time. LLMs accept the same categories and descriptions and generate extraction results. The supplied inputs contain shared category definitions without expected names.

Check their ability to accept new categories separately from whether their extractions follow your instructions.

The notebook constrains Qwen’s output to JSON with the specified categories. Valid JSON can still omit entities or assign incorrect categories.