Please do not copy the URL from the browser for citation. The correct URL is 'https://hdl.handle.net/20.500.12185/708' NCHLT Sepedi POS and Lemma annotated corpus Loading... Files Protocol.SADiLaR.LemmatizationSepedi.Final.2026-03-31.doc (114 KB) Protocol.SADiLaR.PartOfSpeechTaggingSepedi.Final.2026-03-31.docx (55.98 KB) README.SAD-IV.NCHLT_LEMMA-POS_converted.Final.2026-03-31.txt (4.94 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TEST.2026-03-31.nso.txt (241.71 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TRAIN.2026-03-31.nso.txt (2.19 MB) Date 2026-03-31 Authors Gaustad, Tanja Journal Title Journal ISSN Volume Title Publisher North-West University - Centre for Text Technology (CTexT) Abstract Description NCHLT corpora with tokens lemmatised and converted to POS tags used during the SADiLaR-II project for Sepedi. The POS tag conversion results have been thoroughly quality controlled by linguistic experts. Part of the multilingual NCHLT data set where each text type data file contains approximately 45,000 tokens for the conjunctive languages and 75,000 tokens for the disjunctive languages. Sepedi dataset (Lemma, POS annotated) TRAIN: 65,920 tokens; TEST: 7,157 tokens; Total: 73,077 tokens Please see the included protocols for more details on the POS tags used and on the lemmatisation process. The data is given as txt files where each line contains a token, the corresponding lemma, morphological analysis and POS tag, all tab separated. The morphological information originates from a previous project "SADiLaR II (Extension): Linguistic corpus enrichment for South African languages" (see handles below for the data with only morphological analysis included). The data has been split into Train and Test sets according to the original NCHLT data (see handles below for the original NCHLT data). NB: There can be tokenisation differences between this release of the NCHLT data, the original data and the morphologically annotated data due to corrections. Morphological information included from: Morphologically annotated corpus for Sepedi https://hdl.handle.net/20.500.12185/675 Original NCHLT data: NCHLT Sepedi Annotated Text Corpora https://hdl.handle.net/20.500.12185/325 Keywords Sepedi, POS annotated, Lemma annotated, NCHLT, annotated corpus Citation License Creative Commons Attribution 4.0 International URI https://hdl.handle.net/20.500.12185/708 Collections Resource Catalogue Verification status Level 0 Full item page
Please do not copy the URL from the browser for citation. The correct URL is 'https://hdl.handle.net/20.500.12185/708'
NCHLT Sepedi POS and Lemma annotated corpus
Loading... Files Protocol.SADiLaR.LemmatizationSepedi.Final.2026-03-31.doc (114 KB) Protocol.SADiLaR.PartOfSpeechTaggingSepedi.Final.2026-03-31.docx (55.98 KB) README.SAD-IV.NCHLT_LEMMA-POS_converted.Final.2026-03-31.txt (4.94 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TEST.2026-03-31.nso.txt (241.71 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TRAIN.2026-03-31.nso.txt (2.19 MB) Date 2026-03-31 Authors Gaustad, Tanja Journal Title Journal ISSN Volume Title Publisher North-West University - Centre for Text Technology (CTexT) Abstract Description NCHLT corpora with tokens lemmatised and converted to POS tags used during the SADiLaR-II project for Sepedi. The POS tag conversion results have been thoroughly quality controlled by linguistic experts. Part of the multilingual NCHLT data set where each text type data file contains approximately 45,000 tokens for the conjunctive languages and 75,000 tokens for the disjunctive languages. Sepedi dataset (Lemma, POS annotated) TRAIN: 65,920 tokens; TEST: 7,157 tokens; Total: 73,077 tokens Please see the included protocols for more details on the POS tags used and on the lemmatisation process. The data is given as txt files where each line contains a token, the corresponding lemma, morphological analysis and POS tag, all tab separated. The morphological information originates from a previous project "SADiLaR II (Extension): Linguistic corpus enrichment for South African languages" (see handles below for the data with only morphological analysis included). The data has been split into Train and Test sets according to the original NCHLT data (see handles below for the original NCHLT data). NB: There can be tokenisation differences between this release of the NCHLT data, the original data and the morphologically annotated data due to corrections. Morphological information included from: Morphologically annotated corpus for Sepedi https://hdl.handle.net/20.500.12185/675 Original NCHLT data: NCHLT Sepedi Annotated Text Corpora https://hdl.handle.net/20.500.12185/325 Keywords Sepedi, POS annotated, Lemma annotated, NCHLT, annotated corpus Citation License Creative Commons Attribution 4.0 International URI https://hdl.handle.net/20.500.12185/708 Collections Resource Catalogue Verification status Level 0 Full item page
Loading... Files Protocol.SADiLaR.LemmatizationSepedi.Final.2026-03-31.doc (114 KB) Protocol.SADiLaR.PartOfSpeechTaggingSepedi.Final.2026-03-31.docx (55.98 KB) README.SAD-IV.NCHLT_LEMMA-POS_converted.Final.2026-03-31.txt (4.94 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TEST.2026-03-31.nso.txt (241.71 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TRAIN.2026-03-31.nso.txt (2.19 MB) Date 2026-03-31 Authors Gaustad, Tanja Journal Title Journal ISSN Volume Title Publisher North-West University - Centre for Text Technology (CTexT)
Files Protocol.SADiLaR.LemmatizationSepedi.Final.2026-03-31.doc (114 KB) Protocol.SADiLaR.PartOfSpeechTaggingSepedi.Final.2026-03-31.docx (55.98 KB) README.SAD-IV.NCHLT_LEMMA-POS_converted.Final.2026-03-31.txt (4.94 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TEST.2026-03-31.nso.txt (241.71 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TRAIN.2026-03-31.nso.txt (2.19 MB)
Protocol.SADiLaR.LemmatizationSepedi.Final.2026-03-31.doc (114 KB) Protocol.SADiLaR.PartOfSpeechTaggingSepedi.Final.2026-03-31.docx (55.98 KB) README.SAD-IV.NCHLT_LEMMA-POS_converted.Final.2026-03-31.txt (4.94 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TEST.2026-03-31.nso.txt (241.71 KB) SAD-IV.NCHLT_Lemma-Morph-POS.Final_TRAIN.2026-03-31.nso.txt (2.19 MB)
Publisher North-West University - Centre for Text Technology (CTexT)
North-West University - Centre for Text Technology (CTexT)
Abstract Description NCHLT corpora with tokens lemmatised and converted to POS tags used during the SADiLaR-II project for Sepedi. The POS tag conversion results have been thoroughly quality controlled by linguistic experts. Part of the multilingual NCHLT data set where each text type data file contains approximately 45,000 tokens for the conjunctive languages and 75,000 tokens for the disjunctive languages. Sepedi dataset (Lemma, POS annotated) TRAIN: 65,920 tokens; TEST: 7,157 tokens; Total: 73,077 tokens Please see the included protocols for more details on the POS tags used and on the lemmatisation process. The data is given as txt files where each line contains a token, the corresponding lemma, morphological analysis and POS tag, all tab separated. The morphological information originates from a previous project "SADiLaR II (Extension): Linguistic corpus enrichment for South African languages" (see handles below for the data with only morphological analysis included). The data has been split into Train and Test sets according to the original NCHLT data (see handles below for the original NCHLT data). NB: There can be tokenisation differences between this release of the NCHLT data, the original data and the morphologically annotated data due to corrections. Morphological information included from: Morphologically annotated corpus for Sepedi https://hdl.handle.net/20.500.12185/675 Original NCHLT data: NCHLT Sepedi Annotated Text Corpora https://hdl.handle.net/20.500.12185/325 Keywords Sepedi, POS annotated, Lemma annotated, NCHLT, annotated corpus Citation License Creative Commons Attribution 4.0 International URI https://hdl.handle.net/20.500.12185/708 Collections Resource Catalogue Verification status Level 0 Full item page
Description NCHLT corpora with tokens lemmatised and converted to POS tags used during the SADiLaR-II project for Sepedi. The POS tag conversion results have been thoroughly quality controlled by linguistic experts. Part of the multilingual NCHLT data set where each text type data file contains approximately 45,000 tokens for the conjunctive languages and 75,000 tokens for the disjunctive languages. Sepedi dataset (Lemma, POS annotated) TRAIN: 65,920 tokens; TEST: 7,157 tokens; Total: 73,077 tokens Please see the included protocols for more details on the POS tags used and on the lemmatisation process. The data is given as txt files where each line contains a token, the corresponding lemma, morphological analysis and POS tag, all tab separated. The morphological information originates from a previous project "SADiLaR II (Extension): Linguistic corpus enrichment for South African languages" (see handles below for the data with only morphological analysis included). The data has been split into Train and Test sets according to the original NCHLT data (see handles below for the original NCHLT data). NB: There can be tokenisation differences between this release of the NCHLT data, the original data and the morphologically annotated data due to corrections. Morphological information included from: Morphologically annotated corpus for Sepedi https://hdl.handle.net/20.500.12185/675 Original NCHLT data: NCHLT Sepedi Annotated Text Corpora https://hdl.handle.net/20.500.12185/325
NCHLT corpora with tokens lemmatised and converted to POS tags used during the SADiLaR-II project for Sepedi. The POS tag conversion results have been thoroughly quality controlled by linguistic experts. Part of the multilingual NCHLT data set where each text type data file contains approximately 45,000 tokens for the conjunctive languages and 75,000 tokens for the disjunctive languages. Sepedi dataset (Lemma, POS annotated) TRAIN: 65,920 tokens; TEST: 7,157 tokens; Total: 73,077 tokens Please see the included protocols for more details on the POS tags used and on the lemmatisation process. The data is given as txt files where each line contains a token, the corresponding lemma, morphological analysis and POS tag, all tab separated. The morphological information originates from a previous project "SADiLaR II (Extension): Linguistic corpus enrichment for South African languages" (see handles below for the data with only morphological analysis included). The data has been split into Train and Test sets according to the original NCHLT data (see handles below for the original NCHLT data). NB: There can be tokenisation differences between this release of the NCHLT data, the original data and the morphologically annotated data due to corrections. Morphological information included from: Morphologically annotated corpus for Sepedi https://hdl.handle.net/20.500.12185/675 Original NCHLT data: NCHLT Sepedi Annotated Text Corpora https://hdl.handle.net/20.500.12185/325
Keywords Sepedi, POS annotated, Lemma annotated, NCHLT, annotated corpus
Sepedi, POS annotated, Lemma annotated, NCHLT, annotated corpus
License Creative Commons Attribution 4.0 International
Creative Commons Attribution 4.0 International
URI https://hdl.handle.net/20.500.12185/708
https://hdl.handle.net/20.500.12185/708
Collections Resource Catalogue