Digital Humanities B (Language Processing and Information Retrieval)
All Digital Humanities courses
Fall–winter 2026
Class materials
Class 1 · October 6
14:40–16:10 · DH Lab B307. Read the text during class. Lab computers are available.
Course purpose
This course develops practical skills in natural language processing (NLP) and information retrieval for humanities research. Using modern Japanese and English literature, students compare established NLP methods with recent AI methods and evaluate their results, limitations, computational requirements, and reproducibility.
Learning goals
- Prepare and search corpora with metadata.
- Write basic Python text-processing programs and create and interpret visualizations.
- Apply text search, topic models, named entity recognition (NER), and document classification.
- Compare dictionary and rule-based methods, trained models such as GLiNER, and generative large language models (LLMs) using source texts and shared evaluation criteria.
- Produce a reproducible analysis notebook or documented program explaining the research question, method, evidence, and limitations.
Sources and tools
- Natsume Simple (Natsume): Japanese co-occurrence search
- Soranoha: Aozora Bunko texts and bibliographic information
- BERTopic: topic modeling using embeddings
- Gensim / LDA: topic modeling using word frequencies
- Hugging Face Models: model search and downloads
- Hugging Face Datasets: dataset search and downloads
- Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, with Language Models (Daniel Jurafsky and James H. Martin, 2026): 3rd edition; online manuscript released August 19, 2026