Machine Learning

Text-as-Data & NLP Services

Earnings calls, corporate filings, news, central-bank statements, and social media are vast untapped datasets. Text-as-data methods turn that language into measurable variables—sentiment, topics, tone, similarity—that can enter a rigorous analysis. MAS Research builds these measures with modern NLP and validates them so they can be trusted as research data.

Text-as-data methods use natural language processing (NLP) to convert unstructured text into quantitative measures for research—such as sentiment, topics, tone, readability, or similarity. Techniques range from dictionary and topic-modeling approaches to embeddings and large language models. The resulting measures become variables in econometric or statistical analysis, and their validity must be checked like any measurement.

Sentiment, tone & topics Embeddings & LLM classification Filings, calls, news & social Validated, reproducible measures
The text-as-data pipeline Raw text sources processed by NLP into quantitative measures that feed a research analysis. text_as_data · language → measures → analysis text sources filings earnings calls news social media NLP processing measures sentiment topics tone / similarity measures become variables in econometric / statistical analysis validated & used in the model
Text-as-data pipeline text measures

What text-as-data does

A huge amount of economically and managerially meaningful information exists only as text: the tone of an earnings call, the risk language in a 10-K, the stance of a central-bank statement, the sentiment in news coverage or social media. Text-as-data methods treat this language as a data source—using natural language processing to convert it into quantitative measures that can enter a rigorous econometric or statistical analysis, just like any other variable.

The measures come in several families. Sentiment and tone capture how positive, negative, uncertain, or risk-focused a text is—often using finance-specific dictionaries rather than generic ones, because words carry different meaning in financial contexts. Topic modeling discovers the themes running through a corpus and how much each document emphasises each. Embeddings represent words and documents as vectors, enabling similarity, semantic search, and richer downstream models. And increasingly, large language models are used for classification and extraction tasks that older methods handled poorly. The output is a clean, analysable dataset built from language.

When to use it, and the sources we work with

Use text-as-data when the variable you need is latent in documents rather than in a spreadsheet—when a research question hinges on what was said, how, and in what tone. It has become central to finance, accounting, economics, and management research, and to the growing use of alternative data for both research and practice.

We work across the sources these fields rely on: corporate filings (annual reports, 10-Ks, disclosures), earnings-call transcripts, news and media, central-bank communication (statements, minutes, speeches), and social media—as well as building text networks that map relationships between documents, entities, or themes. One principle governs all of it: a text-based measure is a measurement, and like any measurement it has to be validated—checked against human judgement or known benchmarks—before it is trusted as a research variable. An unvalidated sentiment score is not evidence.

At a glance

Approaches to turning text into measures

Methods for building measures from text
ApproachWhat it producesBest for
Dictionary / lexiconSentiment, tone, uncertainty scoresTransparent, replicable measures (e.g. finance dictionaries)
Topic modelingThemes and document-topic weightsDiscovering structure in a large corpus
EmbeddingsVector representations of words/documentsSimilarity, semantic search, downstream models
Supervised classificationLabels (e.g. category, stance)When labelled training data exist
LLM-based classificationLabels / extractions from promptsNuanced tasks older methods handle poorly
Methodology

Building text measures that hold up

The credibility of a text-as-data study rests on treating the text measure as seriously as any other measurement. The decisive step is validation: a sentiment score, topic assignment, or LLM label must be checked against human-coded samples or established benchmarks, with agreement reported—otherwise there is no basis for trusting it. Alongside this, careful pre-processing (cleaning, tokenisation, handling of negation and context) and domain-appropriate tools matter: a general-purpose sentiment dictionary can badly misread financial text, where a word like “liability” is neutral, which is why finance-specific lexicons were developed.

Modern methods bring their own cautions. Topic models require choosing the number of topics and interpreting them judiciously—a statistically fitted topic is not automatically a meaningful theme. LLM-based classification is powerful but can be inconsistent and sensitive to prompt wording, so it needs validation, reproducibility controls, and a check for spurious or hallucinated outputs rather than being trusted at face value. And once a measure is built, the usual rules of inference apply to whatever analysis uses it—including the distinction between prediction and causation. We validate the measure first, then use it properly.

A text measure is a measurement—validate it before you trust it. An unvalidated sentiment score or LLM label is not evidence. We check text-derived measures against human judgement or benchmarks, use domain-appropriate tools (finance text needs finance dictionaries), and report agreement—before the measure enters any model.

Software

We deliver text-as-data in established, reproducible tools—Python (spaCy, NLTK, gensim, scikit-learn, Hugging Face transformers, and LLM APIs) and R (quanteda, tidytext, stm)—with validated measures, documented pre-processing, reproducibility controls for any LLM step, and clean analysis-ready outputs, all with versioned code.

How we work

How we deliver a text-as-data study

Text-as-data sits within our wider machine-learning practice and connects to our econometrics work—so measures are validated, then used in sound analysis.

We start from the construct you need to measure and the text sources available, then design the extraction—choosing dictionary, topic-model, embedding, supervised, or LLM-based methods to fit the task. We assemble and clean the corpus, build the measures, and—before anything else—validate them against human-coded samples or benchmarks, reporting agreement. Only then do the measures enter the downstream analysis.

Reporting sets out the sources and corpus, the extraction method and why it was chosen, the pre-processing, the validation evidence, and the analysis that uses the measures—so the whole pipeline from language to finding is transparent and reproducible.

You receive the validated text-derived measures as an analysis-ready dataset, the validation evidence and documentation of the pipeline, the downstream analysis if we carry it through, and reproducible analytical code and analysis-ready files (where appropriate and permitted)—with any licensing or data-source constraints respected.

Where we apply it

Text-as-data across Management & Allied Studies

Language-rich sources are everywhere in business and economics research—so text-as-data applies across the disciplines we serve.

Finance & Accounting

Earnings-call tone, disclosure and 10-K risk language, and news sentiment as drivers or measures in asset-pricing and corporate research—a leading text-as-data setting.

Economics & Monetary Policy

Central-bank communication analysis—the tone and content of statements, minutes, and speeches—and news-based economic indicators.

Marketing & Consumer Research

Social-media and review analytics—sentiment, themes, and brand perception from large volumes of consumer text.

Management & Strategy

Analysing corporate narratives, strategy disclosures, and communication to measure constructs latent in text.

Political Science & Communication

Analysing speeches, manifestos, media, and social platforms for stance, framing, and topic structure.

Information Systems

Text networks and large-scale content analysis of user-generated and platform data.

FAQ

Text-as-data & NLP: common questions

Text-as-data is the use of natural language processing to turn unstructured text—filings, earnings calls, news, central-bank statements, social media—into quantitative measures such as sentiment, tone, topics, or similarity. Those measures become variables in econometric or statistical analysis. It lets researchers study information that exists only as language, rather than in ready-made numerical datasets.
Often with domain-specific dictionaries rather than general-purpose ones, because words carry different connotations in financial contexts—a general sentiment lexicon can misclassify neutral financial terms. Dictionary methods are transparent and replicable; supervised or LLM-based classification can capture more nuance where labelled data or careful prompting allow. Whichever is used, the resulting measure is validated against human judgement before it is trusted.
Topic modeling discovers the latent themes running through a collection of documents and estimates how much each document emphasises each theme, without pre-specifying the topics. It is useful for mapping the structure of a large corpus. The number of topics must be chosen and the topics interpreted judiciously, because a statistically fitted topic is not automatically a substantively meaningful theme—interpretation and validation matter.
Yes. LLMs can classify and extract information from text for tasks that older methods handle poorly, often with less labelled data. But they can be inconsistent, sensitive to prompt wording, and capable of producing plausible-but-wrong outputs, so LLM-based measures require validation against human-coded samples, reproducibility controls (fixed prompts and settings), and checks for spurious results—the same measurement standard as any other method.
By comparing it against a trusted reference—typically a sample of human-coded documents or an established benchmark—and reporting the level of agreement. A measure that does not track human judgement or a known standard should not be used as evidence. Validation is the step that turns a text-processing output into a credible research variable, and it is central to how we work.
Corporate filings (annual reports, 10-Ks, disclosures), earnings-call transcripts, news and media, central-bank communication (statements, minutes, speeches), social media, and other document collections—as well as building text networks that map relationships between documents, entities, or themes. We work within the licensing and terms of the data sources involved, and respect any access or usage constraints.

Have a research question hiding in text?

Whether it is earnings calls, filings, news, central-bank communication, or social media, we turn the language into validated, analysis-ready measures—and use them in a sound analysis, with the whole pipeline reproducible.