Text-as-Data & NLP Services
Earnings calls, corporate filings, news, central-bank statements, and social media are vast untapped datasets. Text-as-data methods turn that language into measurable variables—sentiment, topics, tone, similarity—that can enter a rigorous analysis. MAS Research builds these measures with modern NLP and validates them so they can be trusted as research data.
Text-as-data methods use natural language processing (NLP) to convert unstructured text into quantitative measures for research—such as sentiment, topics, tone, readability, or similarity. Techniques range from dictionary and topic-modeling approaches to embeddings and large language models. The resulting measures become variables in econometric or statistical analysis, and their validity must be checked like any measurement.
What text-as-data does
A huge amount of economically and managerially meaningful information exists only as text: the tone of an earnings call, the risk language in a 10-K, the stance of a central-bank statement, the sentiment in news coverage or social media. Text-as-data methods treat this language as a data source—using natural language processing to convert it into quantitative measures that can enter a rigorous econometric or statistical analysis, just like any other variable.
The measures come in several families. Sentiment and tone capture how positive, negative, uncertain, or risk-focused a text is—often using finance-specific dictionaries rather than generic ones, because words carry different meaning in financial contexts. Topic modeling discovers the themes running through a corpus and how much each document emphasises each. Embeddings represent words and documents as vectors, enabling similarity, semantic search, and richer downstream models. And increasingly, large language models are used for classification and extraction tasks that older methods handled poorly. The output is a clean, analysable dataset built from language.
When to use it, and the sources we work with
Use text-as-data when the variable you need is latent in documents rather than in a spreadsheet—when a research question hinges on what was said, how, and in what tone. It has become central to finance, accounting, economics, and management research, and to the growing use of alternative data for both research and practice.
We work across the sources these fields rely on: corporate filings (annual reports, 10-Ks, disclosures), earnings-call transcripts, news and media, central-bank communication (statements, minutes, speeches), and social media—as well as building text networks that map relationships between documents, entities, or themes. One principle governs all of it: a text-based measure is a measurement, and like any measurement it has to be validated—checked against human judgement or known benchmarks—before it is trusted as a research variable. An unvalidated sentiment score is not evidence.
Approaches to turning text into measures
| Approach | What it produces | Best for |
|---|---|---|
| Dictionary / lexicon | Sentiment, tone, uncertainty scores | Transparent, replicable measures (e.g. finance dictionaries) |
| Topic modeling | Themes and document-topic weights | Discovering structure in a large corpus |
| Embeddings | Vector representations of words/documents | Similarity, semantic search, downstream models |
| Supervised classification | Labels (e.g. category, stance) | When labelled training data exist |
| LLM-based classification | Labels / extractions from prompts | Nuanced tasks older methods handle poorly |
Building text measures that hold up
The credibility of a text-as-data study rests on treating the text measure as seriously as any other measurement. The decisive step is validation: a sentiment score, topic assignment, or LLM label must be checked against human-coded samples or established benchmarks, with agreement reported—otherwise there is no basis for trusting it. Alongside this, careful pre-processing (cleaning, tokenisation, handling of negation and context) and domain-appropriate tools matter: a general-purpose sentiment dictionary can badly misread financial text, where a word like “liability” is neutral, which is why finance-specific lexicons were developed.
Modern methods bring their own cautions. Topic models require choosing the number of topics and interpreting them judiciously—a statistically fitted topic is not automatically a meaningful theme. LLM-based classification is powerful but can be inconsistent and sensitive to prompt wording, so it needs validation, reproducibility controls, and a check for spurious or hallucinated outputs rather than being trusted at face value. And once a measure is built, the usual rules of inference apply to whatever analysis uses it—including the distinction between prediction and causation. We validate the measure first, then use it properly.
A text measure is a measurement—validate it before you trust it. An unvalidated sentiment score or LLM label is not evidence. We check text-derived measures against human judgement or benchmarks, use domain-appropriate tools (finance text needs finance dictionaries), and report agreement—before the measure enters any model.
Software
We deliver text-as-data in established, reproducible tools—Python (spaCy, NLTK, gensim, scikit-learn, Hugging Face transformers, and LLM APIs) and R (quanteda, tidytext, stm)—with validated measures, documented pre-processing, reproducibility controls for any LLM step, and clean analysis-ready outputs, all with versioned code.
How we deliver a text-as-data study
Text-as-data sits within our wider machine-learning practice and connects to our econometrics work—so measures are validated, then used in sound analysis.
We start from the construct you need to measure and the text sources available, then design the extraction—choosing dictionary, topic-model, embedding, supervised, or LLM-based methods to fit the task. We assemble and clean the corpus, build the measures, and—before anything else—validate them against human-coded samples or benchmarks, reporting agreement. Only then do the measures enter the downstream analysis.
Reporting sets out the sources and corpus, the extraction method and why it was chosen, the pre-processing, the validation evidence, and the analysis that uses the measures—so the whole pipeline from language to finding is transparent and reproducible.
You receive the validated text-derived measures as an analysis-ready dataset, the validation evidence and documentation of the pipeline, the downstream analysis if we carry it through, and reproducible analytical code and analysis-ready files (where appropriate and permitted)—with any licensing or data-source constraints respected.
Text-as-data across Management & Allied Studies
Language-rich sources are everywhere in business and economics research—so text-as-data applies across the disciplines we serve.
Finance & Accounting
Earnings-call tone, disclosure and 10-K risk language, and news sentiment as drivers or measures in asset-pricing and corporate research—a leading text-as-data setting.
Economics & Monetary Policy
Central-bank communication analysis—the tone and content of statements, minutes, and speeches—and news-based economic indicators.
Marketing & Consumer Research
Social-media and review analytics—sentiment, themes, and brand perception from large volumes of consumer text.
Management & Strategy
Analysing corporate narratives, strategy disclosures, and communication to measure constructs latent in text.
Political Science & Communication
Analysing speeches, manifestos, media, and social platforms for stance, framing, and topic structure.
Information Systems
Text networks and large-scale content analysis of user-generated and platform data.
Text-as-data & NLP: common questions
Have a research question hiding in text?
Whether it is earnings calls, filings, news, central-bank communication, or social media, we turn the language into validated, analysis-ready measures—and use them in a sound analysis, with the whole pipeline reproducible.