HOME

TheInfoList



OR:

Corpus linguistics is the study of a language as that language is expressed in its
text corpus In linguistics, a corpus (plural ''corpora'') or text corpus is a language resource consisting of a large and structured set of texts (nowadays usually electronically stored and processed). In corpus linguistics, they are used to do statistical a ...
(plural ''corpora''), its body of "real world" text. Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. The text-corpus method uses the body of texts written in any natural language to derive the set of abstract rules which govern that language. Those results can be used to explore the relationships between that subject language and other languages which have undergone a similar analysis. The first such corpora were manually derived from source texts, but now that work is automated. Corpora have not only been used for linguistics research, they have also been used to compile
dictionaries A dictionary is a listing of lexemes from the lexicon of one or more specific languages, often arranged alphabetically (or by radical and stroke for ideographic languages), which may include information on definitions, usage, etymologies, p ...
(starting with ''
The American Heritage Dictionary of the English Language ''The American Heritage Dictionary of the English Language'' (''AHD'') is an American English, American dictionary of English published by Boston publisher Houghton Mifflin Harcourt, Houghton Mifflin, the first edition of which appeared in 1969. ...
'' in 1969) and grammar guides, such as ''
A Comprehensive Grammar of the English Language ''A Comprehensive Grammar of the English Language'' is a descriptive grammar of English written by Randolph Quirk, Sidney Greenbaum, Geoffrey Leech, and Jan Svartvik. It was first published by Longman in 1985. In 1991 it was called "The greates ...
'', published in 1985. Experts in the field have differing views about the annotation of a corpus. These views range from
John McHardy Sinclair John McHardy Sinclair (14 June 1933 – 13 March 2007) was a Professor of Modern English Language at Birmingham University from 1965 to 2000. He pioneered work in corpus linguistics, discourse analysis, lexicography, and language teaching ...
, who advocates minimal annotation so texts speak for themselves, to the
Survey of English Usage The Survey of English Usage was the first research centre in Europe to carry out research with corpora. The Survey is based in the Department of English Language and Literature at University College London. History The Survey of English Usage wa ...
team (
University College In a number of countries, a university college is a college institution that provides tertiary education but does not have full or independent university status. A university college is often part of a larger university. The precise usage varies ...
, London), who advocate annotation as allowing greater linguistic understanding through rigorous recording.


History

Some of the earliest efforts at grammatical description were based at least in part on corpora of particular religious or cultural significance. For example, Prātiśākhya literature described the sound patterns of
Sanskrit Sanskrit (; attributively , ; nominally , , ) is a classical language belonging to the Indo-Aryan branch of the Indo-European languages. It arose in South Asia after its predecessor languages had diffused there from the northwest in the late ...
as found in the
Vedas upright=1.2, The Vedas are ancient Sanskrit texts of Hinduism. Above: A page from the '' Atharvaveda''. The Vedas (, , ) are a large body of religious texts originating in ancient India. Composed in Vedic Sanskrit, the texts constitute the ...
, and
Pāṇini , era = ;;6th–5th century BCE , region = Indian philosophy , main_interests = Grammar, linguistics , notable_works = ' (Sanskrit#Classical Sanskrit, Classical Sanskrit) , influenced= , notable_ideas=Descript ...
's grammar of
classical Sanskrit Sanskrit (; attributively , ; nominally , , ) is a classical language belonging to the Indo-Aryan branch of the Indo-European languages. It arose in South Asia after its predecessor languages had diffused there from the northwest in the late ...
was based at least in part on analysis of that same corpus. Similarly, the early
Arabic grammarians Arabic grammar or Arabic language sciences ( ar, النحو العربي ' or ar, عُلُوم اللغَة العَرَبِيَّة ') is the grammar of the Arabic language. Arabic is a Semitic language and its grammar has many similarities with ...
paid particular attention to the language of the
Quran The Quran (, ; Standard Arabic: , Classical Arabic, Quranic Arabic: , , 'the recitation'), also romanized Qur'an or Koran, is the central religious text of Islam, believed by Muslims to be a revelation in Islam, revelation from God in Islam, ...
. In the Western European tradition, scholars prepared concordances to allow detailed study of the language of the Bible and other canonical texts.


English corpora

A landmark in modern corpus linguistics was the publication of ''Computational Analysis of Present-Day American English'' in 1967. Written by
Henry Kučera Henry Kučera (15 February 1925 – 20 February 2010), born Jindřich Kučera () was a Czech-American linguist who pioneered corpus linguistics, linguistic software, a major contributor to the ''American Heritage Dictionary'', and a pioneer in ...
and
W. Nelson Francis W. Nelson Francis (October 23, 1910 – June 14, 2002) was an American author, linguist, and university professor. He served as a member of the faculties of Franklin & Marshall College and Brown University, where he specialized in Engl ...
, the work was based on an analysis of the
Brown Corpus The Brown University Standard Corpus of Present-Day American English (or just Brown Corpus) is an electronic collection of text samples of American English, the first major structured corpus of varied genres. This corpus first set the bar for the ...
, which was a contemporary compilation of about a million American English words, carefully selected from a wide variety of sources. Kučera and Francis subjected the Brown Corpus to a variety of computational analyses and then combined elements of linguistics, language teaching,
psychology Psychology is the scientific study of mind and behavior. Psychology includes the study of conscious and unconscious phenomena, including feelings and thoughts. It is an academic discipline of immense scope, crossing the boundaries betwe ...
, statistics, and sociology to create a rich and variegated opus. A further key publication was
Randolph Quirk Charles Randolph Quirk, Baron Quirk, CBE, FBA (12 July 1920 – 20 December 2017) was a British linguist and life peer. He was the Quain Professor of English language and literature at University College London from 1968 to 1981. He sat as ...
's "Towards a description of English Usage" in 1960 in which he introduced the Survey of English Usage. Shortly thereafter, Boston publisher
Houghton-Mifflin Houghton Mifflin Harcourt (; HMH) is an American publisher of textbooks, instructional technology materials, assessments, reference works, and fiction and non-fiction for both young readers and adults. The company is based in the Boston Financ ...
approached Kučera to supply a million-word, three-line citation base for its new ''
American Heritage Dictionary American(s) may refer to: * American, something of, from, or related to the United States of America, commonly known as the "United States" or "America" ** Americans, citizens and nationals of the United States of America ** American ancestry, pe ...
'', the first
dictionary A dictionary is a listing of lexemes from the lexicon of one or more specific languages, often arranged alphabetically (or by radical and stroke for ideographic languages), which may include information on definitions, usage, etymologies ...
compiled using corpus linguistics. The ''AHD'' took the innovative step of combining prescriptive elements (how language ''should'' be used) with descriptive information (how it actually ''is'' used). Other publishers followed suit. The British publisher Collins'
COBUILD COBUILD, an acronym for Collins Birmingham University International Language Database, is a British research facility set up at the University of Birmingham in 1980 and funded by Collins publishers. The facility was initially led by Professor Jo ...
monolingual learner's dictionary A monolingual learner's dictionary (MLD) is designed to meet the reference needs of people learning a foreign language. MLDs are based on the premise that language-learners should progress from a bilingual dictionary to a monolingual one as they b ...
, designed for users learning
English as a foreign language English as a second or foreign language is the use of English by speakers with different native languages. Language education for people learning English may be known as English as a second language (ESL), English as a foreign language (EF ...
, was compiled using the
Bank of English The Bank of English is a representative subset of the 4.5 billion words COBUILD corpus, a collection of English texts. These are mainly British in origin, but content from North America, Australia, New Zealand, South Africa and other Commonwealth ...
. The
Survey of English Usage The Survey of English Usage was the first research centre in Europe to carry out research with corpora. The Survey is based in the Department of English Language and Literature at University College London. History The Survey of English Usage wa ...
Corpus was used in the development of one of the most important Corpus-based Grammars, which was written by Quirk ''et al.'' and published in 1985 as ''A Comprehensive Grammar of the English Language''. The
Brown Corpus The Brown University Standard Corpus of Present-Day American English (or just Brown Corpus) is an electronic collection of text samples of American English, the first major structured corpus of varied genres. This corpus first set the bar for the ...
has also spawned a number of similarly structured corpora: the
LOB Corpus The Lancaster-Oslo/Bergen (LOB) Corpus is a million-word collection of British English texts which was compiled in the 1970s in collaboration between the University of Lancaster, the University of Oslo, and the Norwegian Computing Centre for the ...
(1960s
British English British English (BrE, en-GB, or BE) is, according to Lexico, Oxford Dictionaries, "English language, English as used in Great Britain, as distinct from that used elsewhere". More narrowly, it can refer specifically to the English language in ...
), Kolhapur (
Indian English Indian English (IE) is a group of English dialects spoken in the republic of India and among the Indian diaspora. English is used by the Indian government for communication, along with Hindi, as enshrined in the Constitution of India. E ...
), Wellington (
New Zealand English New is an adjective referring to something recently made, discovered, or created. New or NEW may refer to: Music * New, singer of K-pop group The Boyz Albums and EPs * ''New'' (album), by Paul McCartney, 2013 * ''New'' (EP), by Regurgitator, ...
), Australian Corpus of English (
Australian English Australian English (AusE, AusEng, AuE, AuEng, en-AU) is the set of varieties of the English language native to Australia. It is the country's common language and ''de facto'' national language; while Australia has no official language, Engli ...
), the Frown Corpus (early 1990s
American English American English, sometimes called United States English or U.S. English, is the set of variety (linguistics), varieties of the English language native to the United States. English is the Languages of the United States, most widely spoken lan ...
), and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the
International Corpus of English The International Corpus of English (ICE) is a set of corpora representing varieties of English from around the world. Over twenty countries or groups of countries where English is the first language or an official second language are included. His ...
, and the
British National Corpus The British National Corpus (BNC) is a 100-million-word text corpus of samples of written and spoken English from a wide range of sources. The corpus covers British English of the late 20th century from a wide variety of genres, with the intention ...
, a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities (
Oxford Oxford () is a city in England. It is the county town and only city of Oxfordshire. In 2020, its population was estimated at 151,584. It is north-west of London, south-east of Birmingham and north-east of Bristol. The city is home to the ...
and Lancaster) and the
British Library The British Library is the national library of the United Kingdom and is one of the largest libraries in the world. It is estimated to contain between 170 and 200 million items from many countries. As a legal deposit library, the British ...
. For contemporary American English, work has stalled on the
American National Corpus The American National Corpus (ANC) is a text corpus of American English containing 22 million words of written and spoken data produced since 1990. Currently, the ANC includes a range of genres, including emerging genres such as email, tweets, and ...
, but the 400+ million word
Corpus of Contemporary American English The Corpus of Contemporary American English (COCA) is a one-billion-word corpus of contemporary American English. It was created by Mark Davies, retired professor of corpus linguistics at Brigham Young University (BYU). Content The Corpus of C ...
(1990–present) is now available through a web interface. The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project, containing one million words, which inspired
Shana Poplack Shana Poplack, is a Distinguished University Professor in the linguistics department of the University of Ottawa and three time holder of the Canada Research Chair (Tier I) in Linguistics. She is a leading proponent of variation theory, the appr ...
's much larger corpus of spoken French in the Ottawa-Hull area.


Multilingual Corpora

In the 1990s, many of the notable early successes on statistical methods in natural-language programming (NLP) occurred in the field of
machine translation Machine translation, sometimes referred to by the abbreviation MT (not to be confused with computer-aided translation, machine-aided human translation or interactive translation), is a sub-field of computational linguistics that investigates t ...
, due especially to work at IBM Research. These systems were able to take advantage of existing multilingual textual corpora that had been produced by the
Parliament of Canada The Parliament of Canada (french: Parlement du Canada) is the federal legislature of Canada, seated at Parliament Hill in Ottawa, and is composed of three parts: the King, the Senate, and the House of Commons. By constitutional convention, the ...
and the
European Union The European Union (EU) is a supranational political and economic union of member states that are located primarily in Europe. The union has a total area of and an estimated total population of about 447million. The EU has often been des ...
as a result of laws calling for the translation of all governmental proceedings into all official languages of the corresponding systems of government. There are corpora in non-European languages as well. For example, the National Institute for Japanese Language and Linguistics in Japan has built a number of corpora of spoken and written Japanese.


Ancient languages corpora

Besides these corpora of living languages, computerized corpora have also been made of collections of texts in ancient languages. An example is the
Andersen Andersen () is a Danish-Norwegian patronymic surname meaning "son of Anders" (itself derived from the Greek name " Ανδρέας/Andreas", cf. English Andrew). It is the fifth most common surname in Denmark, shared by about 3.2% of the population.< ...
-Forbes database of the Hebrew Bible, developed since the 1970s, in which every clause is parsed using graphs representing up to seven levels of syntax, and every segment tagged with seven fields of information. The
Quranic Arabic Corpus The Quranic Arabic Corpus is an annotated linguistic resource consisting of 77,430 words of Quranic Arabic. The project aims to provide morphological and syntactic annotations for researchers wanting to study the language of the Quran. K. Dukes, ...
is an annotated corpus for the Classical Arabic language of the
Quran The Quran (, ; Standard Arabic: , Classical Arabic, Quranic Arabic: , , 'the recitation'), also romanized Qur'an or Koran, is the central religious text of Islam, believed by Muslims to be a revelation in Islam, revelation from God in Islam, ...
. This is a recent project with multiple layers of annotation including morphological segmentation,
part-of-speech tagging In corpus linguistics, part-of-speech tagging (POS tagging or PoS tagging or POST), also called grammatical tagging is the process of marking up a word in a text (corpus) as corresponding to a particular part of speech, based on both its definitio ...
, and syntactic analysis using dependency grammar. The Digital Corpus of Sanskrit (DCS) is a "Sandhi-split corpus of Sanskrit texts with full morphological and lexical analysis... designed for text-historical research in Sanskrit linguistics and philology."


Corpora from specific fields

Besides pure linguistic inquiry, researchers had begun to apply corpus linguistics to other academic and professional fields, such as the emerging sub-discipline of Law and Corpus Linguistics, which seeks to understand legal texts using corpus data and tools. The
DBLP DBLP is a computer science bibliography website. Starting in 1993 at Universität Trier in Germany, it grew from a small collection of HTML files and became an organization hosting a database and logic programming bibliography site. Since Nove ...
Discovery Dataset concentrates on
computer science Computer science is the study of computation, automation, and information. Computer science spans theoretical disciplines (such as algorithms, theory of computation, information theory, and automation) to Applied science, practical discipli ...
, containing relevant computer science publications with sentient metadata such as author affiliations, citations, or study fields. A more focused dataset was introduced by NLP Scholar, a combination of papers of the ACL Anthology and
Google Scholar Google Scholar is a freely accessible web search engine that indexes the full text or metadata of scholarly literature across an array of publishing formats and disciplines. Released in beta in November 2004, the Google Scholar index includes p ...
metadata.


Methods

Corpus linguistics has generated a number of research methods, which attempt to trace a path from data to theory. Wallis and Nelson (2001) first introduced what they called the 3A perspective: Annotation, Abstraction and Analysis. * Annotation consists of the application of a scheme to texts. Annotations may include structural markup,
part-of-speech In grammar, a part of speech or part-of-speech (abbreviated as POS or PoS, also known as word class or grammatical category) is a category of words (or, more generally, of lexical items) that have similar grammatical properties. Words that are ass ...
tagging,
parsing Parsing, syntax analysis, or syntactic analysis is the process of analyzing a string of symbols, either in natural language, computer languages or data structures, conforming to the rules of a formal grammar. The term ''parsing'' comes from Lati ...
, and numerous other representations. * Abstraction consists of the translation (mapping) of terms in the scheme to terms in a theoretically motivated model or dataset. Abstraction typically includes linguist-directed search but may include e.g., rule-learning for parsers. * Analysis consists of statistically probing, manipulating and generalising from the dataset. Analysis might include statistical evaluations, optimisation of rule-bases or knowledge discovery methods. Most lexical corpora today are part-of-speech-tagged (POS-tagged). However even corpus linguists who work with 'unannotated plain text' inevitably apply some method to isolate salient terms. In such situations annotation and abstraction are combined in a lexical search. The advantage of publishing an annotated corpus is that other users can then perform experiments on the corpus (through corpus managers). Linguists with other interests and differing perspectives than the originators' can exploit this work. By sharing data, corpus linguists are able to treat the corpus as a locus of linguistic debate and further study.


See also

* ''
A Linguistic Atlas of Early Middle English ''A Linguistic Atlas of Early Middle English'' (LAEME) is a digital, corpus-driven, historical dialect resource for Early Middle English (1150–1325). LAEME combines a searchable Corpus of Tagged Texts (CTT), an Index of Sources, and dot maps ...
'' *
Collocation In corpus linguistics, a collocation is a series of words or terms that co-occur more often than would be expected by chance. In phraseology, a collocation is a type of compositional phraseme, meaning that it can be understood from the words th ...
*
Collostructional analysis Collostructional analysis is a family of methods developed by (in alphabetical order) Stefan Th. Gries (University of California, Santa Barbara) and Anatol Stefanowitsch (Free University of Berlin). Collostructional analysis aims at measuring the ...
*
Concordance Concordance may refer to: * Agreement (linguistics), a form of cross-reference between different parts of a sentence or phrase * Bible concordance, an alphabetical listing of terms in the Bible * Concordant coastline, in geology, where beds, or la ...
( KWIC) * European Language Resource Association *
Keyword (linguistics) In corpus linguistics a key word is a word which occurs in a text more often than we would expect to occur by chance alone.Scott, M. & Tribble, C., 2006, ''Textual Patterns: keyword and corpus analysis in language education'', Amsterdam: Benjamins, ...
*
Linguistic Data Consortium The Linguistic Data Consortium is an open consortium of universities, companies and government research laboratories. It creates, collects and distributes speech and text databases, lexicons, and other resources for linguistics research and developm ...
* List of text corpora *
Machine translation Machine translation, sometimes referred to by the abbreviation MT (not to be confused with computer-aided translation, machine-aided human translation or interactive translation), is a sub-field of computational linguistics that investigates t ...
*
Natural Language Toolkit The Natural Language Toolkit, or more commonly NLTK, is a suite of libraries and programs for symbolic and statistical natural language processing (NLP) for English written in the Python programming language. It was developed by Steven Bird and E ...
* Pattern grammar *
Search engines A search engine is a software system designed to carry out web searches. They search the World Wide Web in a systematic way for particular information specified in a textual web search query. The search results are generally presented in a ...
: they access the "web corpus" *
Semantic prosody Semantic prosody, also discourse prosody, describes the way in which certain seemingly neutral words can be perceived with positive or negative associations through frequent occurrences with particular collocations. Coined in analogy to linguistic ...
*
Speech corpus A speech corpus (or spoken corpus) is a database of speech audio files and text transcriptions. In speech technology, speech corpora are used, among other things, to create acoustic models (which can then be used with a speech recognition or spea ...
*
Text corpus In linguistics, a corpus (plural ''corpora'') or text corpus is a language resource consisting of a large and structured set of texts (nowadays usually electronically stored and processed). In corpus linguistics, they are used to do statistical a ...
*
Translation memory A translation memory (TM) is a database that stores "segments", which can be sentences, paragraphs or sentence-like units (headings, titles or elements in a list) that have previously been translated, in order to aid human translators. The translati ...
*
Treebank In linguistics, a treebank is a parsed text corpus that annotates syntactic or semantic sentence structure. The construction of parsed corpora in the early 1990s revolutionized computational linguistics, which benefitted from large-scale empiri ...
*
Word list A word list (or ''lexicon'') is a list of a language's lexicon (generally sorted by frequency of occurrence either by levels or as a ranked list) within some given text corpus, serving the purpose of vocabulary acquisition. A lexicon sorted by ...


Notes and references


Further reading


Books

* Biber, D., Conrad, S., Reppen R. ''Corpus Linguistics, Investigating Language Structure and Use'', Cambridge: Cambridge UP, 1998. * McCarthy, D., and Sampson G. ''Corpus Linguistics: Readings in a Widening Discipline'', Continuum, 2005. * Facchinetti, R. ''Theoretical Description and Practical Applications of Linguistic Corpora''. Verona: QuiEdit, 2007 * Facchinetti, R. (ed.) ''Corpus Linguistics 25 Years on''. New York/Amsterdam: Rodopi, 2007 * Facchinetti, R. and Rissanen M. (eds.) ''Corpus-based Studies of Diachronic English''. Bern: Peter Lang, 2006 * Lenders, W. ''Computational lexicography and corpus linguistics until ca. 1970/1980'', in: Gouws, R. H., Heid, U., Schweickard, W., Wiegand, H. E. (eds.) ''Dictionaries – An International Encyclopedia of Lexicography. Supplementary Volume: Recent Developments with Focus on Electronic and Computational Lexicography''. Berlin: De Gruyter Mouton, 2013 * Fuß, Eric et al. (Eds.): ''Grammar and Corpora 2016'', Heidelberg: Heidelberg University Publishing, 2018.
digital open access
. * Stefanowitsch A. 2020. ''Corpus linguistics: A guide to the methodology''. Berlin: Language Science Press. , Open Access https://langsci-press.org/catalog/book/148.


Book series

Book series in this field include: * Language and Computers (Brill)
Studies in Corpus Linguistics (John Benjamins)

English Corpus Linguistics (Peter Lang)
*
Corpus and Discourse (Bloomsbury)


Journals

There are several international peer-reviewed journals dedicated to corpus linguistics, for example: *
Corpora Corpus is Latin for "body". It may refer to: Linguistics * Text corpus, in linguistics, a large and structured set of texts * Speech corpus, in linguistics, a large set of speech audio files * Corpus linguistics, a branch of linguistics Music * ...
* Corpus Linguistics and Linguistic Theory
ICAME Journal
*
International Journal of Corpus Linguistics The ''International Journal of Corpus Linguistics'' is a quarterly peer-reviewed academic journal that publishes scholarly articles and book reviews on corpus linguistics, with a focus on applied linguistics. The journal is published by John Benjam ...

Language Resources and Evaluation Journal
supported by th
European Language Resources Association

Research in Corpus Linguistics
supported by the Spanish Association for Corpus Linguistics (AELINCO)


External links



* ttps://web.archive.org/web/20060113235630/http://torvald.aksis.uib.no/corpora/ Corpora discussion list
Freely-available, web-based corpora (100 million – 400 million words each): American (COCA, COHA), British (BNC), ''Time'', Spanish, Portuguese



Przemek Kaszubski's list of references

AskOxford.com
''the composition and use of the Oxford Corpus''
DMCBC.com

Datum Multilanguage Corpora Based on chinese free sample download

Corpus4u Community
a Chinese online forum for corpus linguistics
McEnery and Wilson's Corpus Linguistics Page

Corpus Linguistics with R mailing list

Research and Development Unit for English Studies

Survey of English Usage

The Centre for Corpus Linguistics at Birmingham University

Tools for Corpus Linguistics (annotated list)

Gateway to Corpus Linguistics on the Internet
an annotated guide to corpus resources on the web
Biomedical corpora

Linguistic Data Consortium
a major distributor of corpora
Penn Parsed Corpora of Historical English

Corsis
(formerly Tenka Text) an
open-source Open source is source code that is made freely available for possible modification and redistribution. Products include permission to use the source code, design documents, or content of the product. The open-source model is a decentralized sof ...
(
GPL The GNU General Public License (GNU GPL or simply GPL) is a series of widely used free software licenses that guarantee end users the four freedoms to run, study, share, and modify the software. The license was the first copyleft for general u ...
ed) corpus analysis tool written in C#
ICECUP
an
Fuzzy Tree Fragments

Discussion group
text mining Text mining, also referred to as ''text data mining'', similar to text analytics, is the process of deriving high-quality information from text. It involves "the discovery by computer of new, previously unknown information, by automatically extract ...
* A corpus linguistics related conference MAG 2017: You can find some information and events related t
Metadiscourse Across Genres by visiting MAG 2017 website

Corpus of Political Speeches
Free access to political speeches by American and Chinese politicians, developed by Hong Kong Baptist University Library
LightTag -Text Annotation Tool
A text annotation tool for machine learning corpus focused on team management * LIVAC Synchronous Corpus {{DEFAULTSORT:Corpus Linguistics Applied linguistics Discourse analysis Linguistic history Linguistic research