Annotate text corpora, train classifiers,
without writing code.

ActiveTigger is an open-source web tool for collaborative corpus annotation and classification. Built by and for computational social scientists, it handles text as well as images.

Upload a corpus, define your categories, annotate — and let active learning, fine-tuned encoder models, or generative AI do the heavy lifting when you want them.

It is constantly developed by researchers at CSS@IP-Paris (CREST) and used daily in ongoing research projects.

We host an online instance at CREST — researchers can request an account to try it.

1. Explore your corpus

Get a first sense of your data: browse the texts, search by keywords, visualize the corpus as a 2D projection, or run topic models to see what it contains before deciding how to annotate it.
ActiveTigger exploration view of a corpus, with search and visualisation tools

2. Write your codebook

Define your annotation scheme and write down the guidelines for each category, so they stay consistent across annotators and over time.
ActiveTigger codebook view with annotation guidelines and label counts

3. Annotate together

Tag texts one by one, filter by content or by tag, and let active learning pick the next most informative example. Multiple annotators can work in parallel and compare results.
ActiveTigger annotation interface, with a text to tag and label buttons

4. Train & predict

Fine-tune a BERT classifier on your annotations, evaluate its performance, and apply it to the whole corpus to classify everything you didn't annotate by hand.
ActiveTigger prediction view, with a fine-tuned model applied to the corpus

5. Export everything

Download your annotations, predictions, fine-tuned models, and a full project summary in open formats, ready for analysis in Python or R — or for sharing with your colleagues.
ActiveTigger export view listing downloadable annotations, predictions and models

Your corpus, your categories, your classifier

With ActiveTigger, scholars across the world develop and stabilize codebooks, build training datasets, classify large corpora, and extract information. It is designed to facilitate iterative, reflexive coding. It has been tested on corpora of hundreds of thousands of documents.

Your language

Multilingual models are supported.

Your privacy settings

Because ActiveTigger is open source, you can run your own instance on your own infrastructure when working with sensitive data. No vendor, no data harvesting, no lock-in.

About

ActiveTigger is developed at the CREST research unit (CNRS, Évole Polytechnique, ENSAE) by the CSS@IP-Paris team, with the support of OuestWare. It is a complete rewrite of the original R Shiny app by Julien Boelaert and Étienne Ollion. The name is a pun on the similarity between "Tagger" and "Tiger".

We aim to keep the application lightweight, resource-efficient and open source, and to make computational social science methods accessible to communities beyond programmers. The project was supported by DRARI Île-de-France, Progedo, CREST.

Citing ActiveTigger

If you use ActiveTigger in your research, please cite the companion article (available in the JADT 2026 proceedings):

Schultz, E., Boelaert, J., Morin, A., Bonutti D'Agostini, E., Claesson, A., Chatelain, A., & Ollion, É. (2026). ActiveTigger: An open source collaborative text annotation software for computational social sciences. In Proceedings of the 18th International Conference on Statistical Analysis of Textual Data (JADT 2026), Palermo, Italy, July 8–10, 2026.