#  Textual Analysis 

 



 ##  

  expand\_more  

 
  

 

### **Web-based Tools and Corpora:**

[**Voyant Tools**](http://voyant-tools.org/?lang=ja)

This is the Japanese version of the web-based software, Voyant Tools. It can tokenize Japanese text and eliminates the need to insert whitespace before pasting or uploading text.

[**KH Coder**](http://khcoder.net/en/)

An open source software for quantitative content analysis and text mining. It supports Japanese, English, and numerous other languages. An [English reference manual](http://khcoder.net/en/manual_en_v3.pdf) is available.

[**Center for Open Data in the Humanities**](http://codh.rois.ac.jp/) **(CODH)**

Since 2017 the CODH has been focused on opening up new possibilities in Japanese digital humanities and provide access to multiple open datasets, such as: [Pre-Modern Japanese Text](http://codh.rois.ac.jp/pmjt/) (from the National Institute of Japanese Literature), [Dataset of Edo Cooking Recipes,](http://codh.rois.ac.jp/edo-cooking/) [Bunkan Complete Collection](http://codh.rois.ac.jp/bukan/) (biographical and geospatial data related to daimyo and shognate officials), and [Dataset of Modern Magazines](http://codh.rois.ac.jp/modern-magazine/).

[**Center for Corpus Development, NINJAL**](http://pj.ninjal.ac.jp/corpus_center/en/)

Various web-based tools and corpora are available through the National Institute of Japanese Language and Linguisitics (NINJAL). Notable corpus include: [Shonagon](http://www.kotonoha.gr.jp/shonagon/) (Corpus of Contemporary Written Japanese), [Chunagon](https://chunagon.ninjal.ac.jp/auth/login?service=https://chunagon.ninjal.ac.jp/j_spring_cas_security_check) (Corpus of Contemporary Written and Spoken Japanese; with free registration), as well as the [Oxford-NINJAL Corpus of Old Japanese](http://oncoj.ninjal.ac.jp/).

[**JAPANESE.GR.JP (JGJ)**](http://japanese.gr.jp/index.html)

A text analysis project for Japanese linguistic and literary classics, including software, results, and data sets. There is a special focus on waka poetry.

[**Kokalog**](http://kokalog.net/)

A system that enables full text searches of Diet proceedings, as well as a timeline indicating when these terms were used with greatest frequency. The use of double quotes around each search term is advised ("X").

[**Japanese Text Initiative**](http://jti.lib.virginia.edu/japanese/)

A collaborative effort between the University of Virginia Library Electronic Text Center and the University of Pittsburgh East Asian Library to make texts of classical Japanese literature available on the internet.

[**HathiTrust Research Center**](https://analytics.hathitrust.org/)

Supports large-scale computational analysis of the works in the [HathiTrust Digital Library](https://www.hathitrust.org/). Support of Japanese is, however, not well documented.



 

##  Featured Projects 

**[Aozora Search](https://textual-optics-lab.uchicago.edu/aozora)** *Hoyt Long, University of Chicago* An enhanced search interface of the [Aozora Bunko Digital Library](https://www.aozora.gr.jp/), an open source collection of over 12,000 modern Japanese literary texts, which enables users to perform complex keyword searches and filtering. **[How to Use Aozora Search](https://www.youtube.com/watch?v=ovqrj43-M-Q)** [   ![Aozora Search](/sites/g/files/omnuum10046/files/styles/hwp_1_1__720x720_scale/public/jdrc/files/screen_shot_2018-06-13_at_4.58.10_pm.png?itok=cEaeD0i1) 

 ](http://artflsrv02.uchicago.edu/philologic4/aozora/)

 

##  Text Analysis Guides 

**[Japanese Text Mining Guide ](https://scholarblogs.emory.edu/japanese-text-mining/)** *Mark Ravina, Emory University* Includes numerous guides for analyzing Japanese text using R. [   ![Japanese Text Mining](/sites/g/files/omnuum10046/files/styles/hwp_1_1__960x960_scale/public/jdrc/files/screen_shot_2018-06-14_at_3.16.59_pm.png?itok=6e9RTN8r) 

 ](https://scholarblogs.emory.edu/japanese-text-mining/) \_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_

 **[Tutorials and Reviews](http://dhjapan.org/wiki/doku.php?id=tutorials)** Multiple resources are available through the [Digital Humanities Japan Resource Wiki](http://dhjapan.org/wiki/doku.php?id=start). This page provides materials for cleaning data, encoding, OCR, text segmentation, text mining and web scraping.  


 

##  Word Segmenters 

**[Web Chamame](http://chamame.ninjal.ac.jp/)**  Web-based tool hosted by the National Institute for Japanese Language and Linguistics, NINJAL **[KyTea](http://www.phontron.com/kytea/)** Kyoto Text Analysis Toolkit (KyTea) **[Chasen](http://chasen-legacy.osdn.jp/)  Created by the Nara Institute of Science and Technology