Corpus linguistics is the study of a language as that language is expressed in its text corpus (plural ''corpora''), its body of "real world" text. Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. The text-corpus method uses the body of texts written in any natural language to derive the set of abstract rules which govern that language. Those results can be used to explore the relationships between that subject language and other languages which have undergone a similar analysis. The first such corpora were manually derived from source texts, but now that work is automated. Corpora have not only been used for linguistics research, they have also been used to compile dictionaries (starting with '' The American Heritage Dictionary of the English Language'' in 1969) and grammar guides, such as '' A Comprehensive Grammar of the English Language'', published in 1985. Experts in the field have differing views about the annotation of a corpus. These views range from John McHardy Sinclair, who advocates minimal annotation so texts speak for themselves, to the Survey of English Usage team ( University College, London), who advocate annotation as allowing greater linguistic understanding through rigorous recording.

History

Some of the earliest efforts at grammatical description were based at least in part on corpora of particular religious or cultural significance. For example, Prātiśākhya literature described the sound patterns of

Sanskrit Sanskrit (; attributively , ; nominally , , ) is a classical language belonging to the Indo-Aryan languages, Indo-Aryan branch of the Indo-European languages. It arose in South Asia after its predecessor languages had Trans-cultural diffusion ...

as found in the

Vedas upright=1.2, The Vedas are ancient Sanskrit texts of Hinduism. Above: A page from the '' Atharvaveda''. The Vedas (, , ) are a large body of religious texts originating in ancient India. Composed in Vedic Sanskrit, the texts constitute th ...

, and

Pāṇini , era = ;;6th–5th century BCE , region = Indian philosophy , main_interests = Grammar, linguistics , notable_works = ' ( Classical Sanskrit) , influenced= , notable_ideas= Descriptive linguistics (Devana ...

's grammar of classical Sanskrit was based at least in part on analysis of that same corpus. Similarly, the early

Arabic grammarians Arabic grammar or Arabic language sciences ( ar, النحو العربي ' or ar, عُلُوم اللغَة العَرَبِيَّة ') is the grammar of the Arabic language. Arabic is a Semitic language and its grammar has many similarities wit ...

paid particular attention to the language of the

Quran The Quran (, ; Standard Arabic: , Quranic Arabic: , , 'the recitation'), also romanized Qur'an or Koran, is the central religious text of Islam, believed by Muslims to be a revelation from God. It is organized in 114 chapters (pl.: , ...

. In the Western European tradition, scholars prepared concordances to allow detailed study of the language of the Bible and other canonical texts.

English corpora

A landmark in modern corpus linguistics was the publication of ''Computational Analysis of Present-Day American English'' in 1967. Written by Henry Kučera and W. Nelson Francis, the work was based on an analysis of the Brown Corpus, which was a contemporary compilation of about a million American English words, carefully selected from a wide variety of sources. Kučera and Francis subjected the Brown Corpus to a variety of computational analyses and then combined elements of linguistics, language teaching,

psychology Psychology is the science, scientific study of mind and behavior. Psychology includes the study of consciousness, conscious and Unconscious mind, unconscious phenomena, including feelings and thoughts. It is an academic discipline of immens ...

, statistics, and sociology to create a rich and variegated opus. A further key publication was Randolph Quirk's "Towards a description of English Usage" in 1960 in which he introduced the Survey of English Usage. Shortly thereafter, Boston publisher

Houghton-Mifflin Houghton Mifflin Harcourt (; HMH) is an American publisher of textbooks, instructional technology materials, assessments, reference works, and fiction and non-fiction for both young readers and adults. The company is based in the Boston Fina ...

approached Kučera to supply a million-word, three-line citation base for its new '' American Heritage Dictionary'', the first

dictionary A dictionary is a listing of lexemes from the lexicon of one or more specific languages, often arranged alphabetically (or by radical and stroke for ideographic languages), which may include information on definitions, usage, etymologie ...

compiled using corpus linguistics. The ''AHD'' took the innovative step of combining prescriptive elements (how language ''should'' be used) with descriptive information (how it actually ''is'' used). Other publishers followed suit. The British publisher Collins' COBUILD monolingual learner's dictionary, designed for users learning English as a foreign language, was compiled using the Bank of English. The Survey of English Usage Corpus was used in the development of one of the most important Corpus-based Grammars, which was written by Quirk ''et al.'' and published in 1985 as ''A Comprehensive Grammar of the English Language''. The Brown Corpus has also spawned a number of similarly structured corpora: the LOB Corpus (1960s

British English British English (BrE, en-GB, or BE) is, according to Oxford Dictionaries, "English as used in Great Britain, as distinct from that used elsewhere". More narrowly, it can refer specifically to the English language in England, or, more broadl ...

), Kolhapur ( Indian English), Wellington ( New Zealand English), Australian Corpus of English ( Australian English), the Frown Corpus (early 1990s

American English American English, sometimes called United States English or U.S. English, is the set of varieties of the English language native to the United States. English is the most widely spoken language in the United States and in most circumstances ...

), and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the International Corpus of English, and the British National Corpus, a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities (

Oxford Oxford () is a city in England. It is the county town and only city of Oxfordshire. In 2020, its population was estimated at 151,584. It is north-west of London, south-east of Birmingham and north-east of Bristol. The city is home to the ...

and Lancaster) and the

British Library The British Library is the national library of the United Kingdom and is one of the largest libraries in the world. It is estimated to contain between 170 and 200 million items from many countries. As a legal deposit library, the Briti ...

. For contemporary American English, work has stalled on the American National Corpus, but the 400+ million word Corpus of Contemporary American English (1990–present) is now available through a web interface. The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project, containing one million words, which inspired

Shana Poplack Shana Poplack, is a Distinguished University Professor in the linguistics department of the University of Ottawa and three time holder of the Canada Research Chair (Tier I) in Linguistics. She is a leading proponent of variation theory, the appr ...

's much larger corpus of spoken French in the Ottawa-Hull area.

Multilingual Corpora

In the 1990s, many of the notable early successes on statistical methods in natural-language programming (NLP) occurred in the field of

machine translation Machine translation, sometimes referred to by the abbreviation MT (not to be confused with computer-aided translation, machine-aided human translation or interactive translation), is a sub-field of computational linguistics that investigates ...

, due especially to work at IBM Research. These systems were able to take advantage of existing multilingual textual corpora that had been produced by the Parliament of Canada and the

European Union The European Union (EU) is a supranational union, supranational political union, political and economic union of Member state of the European Union, member states that are located primarily in Europe, Europe. The union has a total area of ...

as a result of laws calling for the translation of all governmental proceedings into all official languages of the corresponding systems of government. There are corpora in non-European languages as well. For example, the National Institute for Japanese Language and Linguistics in Japan has built a number of corpora of spoken and written Japanese.

Ancient languages corpora

Besides these corpora of living languages, computerized corpora have also been made of collections of texts in ancient languages. An example is the Andersen-Forbes database of the Hebrew Bible, developed since the 1970s, in which every clause is parsed using graphs representing up to seven levels of syntax, and every segment tagged with seven fields of information. The Quranic Arabic Corpus is an annotated corpus for the Classical Arabic language of the

. This is a recent project with multiple layers of annotation including morphological segmentation, part-of-speech tagging, and syntactic analysis using dependency grammar. The Digital Corpus of Sanskrit (DCS) is a "Sandhi-split corpus of Sanskrit texts with full morphological and lexical analysis... designed for text-historical research in Sanskrit linguistics and philology."

Corpora from specific fields

Besides pure linguistic inquiry, researchers had begun to apply corpus linguistics to other academic and professional fields, such as the emerging sub-discipline of Law and Corpus Linguistics, which seeks to understand legal texts using corpus data and tools. The DBLP Discovery Dataset concentrates on

computer science Computer science is the study of computation, automation, and information. Computer science spans theoretical disciplines (such as algorithms, theory of computation, information theory, and automation) to Applied science, practical discipli ...

, containing relevant computer science publications with sentient metadata such as author affiliations, citations, or study fields. A more focused dataset was introduced by NLP Scholar, a combination of papers of the ACL Anthology and

Google Scholar Google Scholar is a freely accessible web search engine that indexes the full text or metadata of scholarly literature across an array of publishing formats and disciplines. Released in beta in November 2004, the Google Scholar index includes ...

metadata.

Methods

Corpus linguistics has generated a number of research methods, which attempt to trace a path from data to theory. Wallis and Nelson (2001) first introduced what they called the 3A perspective: Annotation, Abstraction and Analysis. * Annotation consists of the application of a scheme to texts. Annotations may include structural markup, part-of-speech tagging,

parsing Parsing, syntax analysis, or syntactic analysis is the process of analyzing a string of symbols, either in natural language, computer languages or data structures, conforming to the rules of a formal grammar. The term ''parsing'' comes from ...

, and numerous other representations. * Abstraction consists of the translation (mapping) of terms in the scheme to terms in a theoretically motivated model or dataset. Abstraction typically includes linguist-directed search but may include e.g., rule-learning for parsers. * Analysis consists of statistically probing, manipulating and generalising from the dataset. Analysis might include statistical evaluations, optimisation of rule-bases or knowledge discovery methods. Most lexical corpora today are part-of-speech-tagged (POS-tagged). However even corpus linguists who work with 'unannotated plain text' inevitably apply some method to isolate salient terms. In such situations annotation and abstraction are combined in a lexical search. The advantage of publishing an annotated corpus is that other users can then perform experiments on the corpus (through corpus managers). Linguists with other interests and differing perspectives than the originators' can exploit this work. By sharing data, corpus linguists are able to treat the corpus as a locus of linguistic debate and further study.

Notes and references

External links

* ttps://web.archive.org/web/20060113235630/http://torvald.aksis.uib.no/corpora/ Corpora discussion list
Freely-available, web-based corpora (100 million – 400 million words each): American (COCA, COHA), British (BNC), ''Time'', Spanish, Portuguese

Przemek Kaszubski's list of references

AskOxford.com
''the composition and use of the Oxford Corpus''
DMCBC.com

Datum Multilanguage Corpora Based on chinese free sample download

Corpus4u Community
a Chinese online forum for corpus linguistics
McEnery and Wilson's Corpus Linguistics Page

Corpus Linguistics with R mailing list

Research and Development Unit for English Studies

Survey of English Usage

The Centre for Corpus Linguistics at Birmingham University

Tools for Corpus Linguistics (annotated list)

Gateway to Corpus Linguistics on the Internet
an annotated guide to corpus resources on the web
Biomedical corpora

Linguistic Data Consortium
a major distributor of corpora
Penn Parsed Corpora of Historical English

Corsis
(formerly Tenka Text) an open-source ( GPLed) corpus analysis tool written in C#
ICECUP
an
Fuzzy Tree Fragments

Discussion group
text mining * A corpus linguistics related conference MAG 2017: You can find some information and events related t
Metadiscourse Across Genres by visiting MAG 2017 website

Corpus of Political Speeches
Free access to political speeches by American and Chinese politicians, developed by Hong Kong Baptist University Library
LightTag -Text Annotation Tool
A text annotation tool for machine learning corpus focused on team management * LIVAC Synchronous Corpus {{DEFAULTSORT:Corpus Linguistics Applied linguistics Discourse analysis Linguistic history Linguistic research

History

English corpora

Multilingual Corpora

Ancient languages corpora

Corpora from specific fields

Methods

See also

Notes and references

Further reading

Books

Book series

Journals

External links