Skip to content

Module Overview

REX extensions are built from modules. Each module is a self-contained piece that collects a specific type of data or adds a specific capability. You choose which modules to include based on what your study needs to measure.

This page describes each module in plain language. For dependencies, configuration keys, and record names, follow the Module Map link on each entry.


What it collects: Nothing directly. rex-core is the foundation every REX extension is built on. It manages the extension lifecycle, handles participant identification, fetches configuration from your server, and provides the communication layer that all other modules use.

Good for: Everything. You cannot build a REX extension without it.

Limitations: Not configurable by researchers directly. You include it and it works in the background.

A few other modules also work behind the scenes and are never chosen on their own: rex-types, rex-data-collection, and openredaction-in-browser.

Technical reference: See Module Map


What it collects: Browsing history via Chrome’s built-in history API. Each record includes the URL visited, page title, timestamp of the visit, and how the participant got there (typed the address, clicked a link, followed a redirect, etc.).

Good for: Studies that need a log of which sites participants visited and when. Good for media diet research, information behavior studies, and any study where the sequence of sites visited matters.

Limitations: The Chrome history API provides a snapshot, not a live stream. It records that a visit happened but not fine-grained behavior during the visit. It also cannot distinguish between a quick accidental click and an intentional extended visit. For dwell time and real-time behavior, use rex-page-events instead or alongside.

Technical reference: See Module Map


What it collects: The links between visits that browser history leaves out: the redirect steps between a click and the page it lands on (for example, through a search ad), and which page opened a new tab.

Good for: Studies that need to know how participants got from one page to another, such as how often search results or social media posts lead to news sites.

Limitations: Only records links from the day the extension is installed, not before. By default it records ID numbers that are joined to rex-history records, not URLs, so it is used together with rex-history.

Technical reference: See Module Map


What it collects: Real-time page lifecycle events: when a page loads, when the participant switches to a different tab, when the browser window gains or loses focus, and how long the participant spent on each page (dwell time). This is more granular than rex-history because it captures events as they happen rather than reading from the history log.

Good for: Studies where attention and engagement matter: how long participants actually spent reading something, whether they were multitasking, which tabs they had open simultaneously. Also useful for detecting whether participants are actively using the browser versus leaving it open in the background.

Limitations: Dwell time is based on tab focus, not eye tracking or scroll position. A participant can “be on” a page without reading it. Does not capture what was on the page.

Technical reference: See Module Map


What it collects: Nothing on its own. rex-lists manages domain allowlists and blocklists that control which sites other modules record. For example, you could configure rex-history to only record visits to a specific list of news sites, or to skip visits to banking sites.

Good for: Any study where you need to limit data collection to particular domains, or protect participant privacy by excluding sensitive categories of sites.

Limitations: The researcher defines the study’s lists in the configuration. Participants cannot change those entries, though with rex-lists-front-end they can add entries of their own.

Technical reference: See Module Map


What it collects: Nothing. This module adds a page to the extension that shows participants which lists are active, and lets them add, edit, and delete entries of their own.

Good for: Studies where transparency and participant control matter for trust or IRB requirements. Participants can see what categories of sites are being monitored, and add sites they do not want recorded.

Limitations: Entries from the study configuration are read-only for participants. You can let participants delete them, but they return at the next sync.

Technical reference: See Module Map


What it collects: Search engine results pages: the list of results participants see when they search on Google, Bing, or DuckDuckGo. Captures the ranked list of results, including titles, URLs, and snippets, and optionally AI-generated answers (such as Google AI Overviews) and news boxes.

Good for: Studies about algorithmic influence on information access, filter bubbles, search behavior research, or any study where the results participants are exposed to matters as much as what they click on.

Limitations: Captures what was shown, not whether the participant read or clicked the results (rex-page-events and rex-history can fill in that gap). Only the three search engines above are supported.

Technical reference: See Module Map


What it collects: Question-and-answer pairs from AI chatbots (ChatGPT, Gemini, and Perplexity) as they happen, in real time as the participant types and receives responses. It can also capture headlines from news homepages and discovery feeds.

Good for: Studies about how people use AI assistants, what questions they ask, how they interact with AI-generated content, or the role of AI tools in information-seeking behavior. The news capture suits studies of news exposure.

Limitations: Captures the conversation as it occurs in the browser. Does not retrieve past conversations (use a spider for that). Only the platforms and news sites it supports are captured.

Technical reference: See Module Map


Spiders: rex-spider plus a platform plugin

Section titled “Spiders: rex-spider plus a platform plugin”

What it collects: Conversations a participant already has on an AI platform. Unlike rex-live-mirror, which captures conversations as they happen, a spider periodically fetches the participant’s saved conversations from the platform, going back as far as the study’s lookback window allows. You choose how much of each conversation to collect: from start and end times only, up to titles and full text.

Platforms: ChatGPT (rex-spider-chatgpt), Gemini (rex-spider-gemini), Google Search AI Mode (rex-spider-google-ai), and Perplexity (rex-spider-perplexity).

Good for: Studies that need historical AI usage data, not just what happens during the study period. Useful when participants may have been using these tools before enrolling. The levels of detail let you match collection to what your IRB protocol allows.

Limitations: Depends on what the platform exposes. If a platform changes how it works, the spider may stop working until it is updated. Requires the participant to be logged in to the platform.

Technical reference: See Module Map

Data sample: See rex-spider-chatgpt data reference for annotated examples of what conversation records look like.


What it collects: Nothing new. It stores records produced by all other active modules locally in the browser, and provides a download button in the extension so participants or researchers can export the data as a JSON file.

Good for: Pilots, small studies, data donation studies, or situations where setting up a server is not practical. Good for evaluating REX, which is why the demo uses it.

Limitations: Data lives only on the participant’s device until they download it. If a participant uninstalls the extension before downloading, the data is lost. Not suitable for large studies where you need reliable data delivery.

Technical reference: See Module Map


What it collects: Nothing new. It transmits records produced by all other active modules to a PDK server.

Good for: Any study beyond the pilot stage. Data flows automatically to your server as participants browse, with no action required from them. Survivable if participants lose their device or uninstall the extension, since data is already on the server.

Limitations: Requires a PDK server to be set up and maintained. See the Local Backend section for how to run one with Docker Desktop.

Technical reference: See Module Map


What it collects: Nothing new. It shows participants a searchable table of the browsing history the extension has collected on their device, with a summary by site, and lets them delete records before sharing.

Good for: Data donation studies, where participants review their data before sending it, and studies where showing participants their own data is part of the design. Based on Web Historian.

Limitations: Works only with rex-history and rex-local-download, since it reads the history records stored on the device.

Technical reference: See Module Map


What it collects: Browser configuration details, currently the participant’s default search engine.

Good for: Studies where the participant’s browser setup is a relevant variable: for example, research on search engine market share, default effects, or studies where knowing whether participants use Google versus Bing versus DuckDuckGo matters for interpretation.

Limitations: Limited to what Chrome exposes through its extension APIs. Cannot detect all browser settings.

Technical reference: See Module Map


What it collects: Nothing new. It processes records produced by other modules and redacts personal information (such as names, email addresses, and phone numbers) before those records are stored or transmitted.

Good for: Studies where participant privacy protection is a priority, or where IRB requirements call for minimizing exposure of personal data. Useful when rex-history or rex-page-events might capture URLs containing query parameters or other identifiable information, and for chatbot conversation text and titles.

Limitations: Redaction is rule-based and may not catch every form of sensitive data. Not a substitute for a full privacy impact assessment.

Technical reference: See Module Map


What it collects: Records when an intervention was applied to a page.

What it does: Changes what participants see while browsing, to apply experimental conditions: blocking or redirecting sites, obscuring pages, and hiding or restyling parts of a page.

Good for: Experimental studies where the researcher wants to manipulate what participants see as an independent variable. Enables A/B testing and behavioral interventions delivered through the browser.

Limitations: Manipulation is applied through the browser extension, so it only affects the participant’s view, not the actual web page. Requires careful ethical consideration and IRB review.

Technical reference: See Module Map


What it collects: Nothing.

What it does: Adds or changes HTTP request headers, and fields in form submissions, on sites that match configured patterns.

Good for: Studies involving network-level interventions that depend on what the browser sends to websites.

Limitations: Requires a solid understanding of HTTP to use safely. Misconfiguration can break websites for participants. Not commonly needed for observational studies.

Technical reference: See Module Map


What it collects: Nothing.

What it does: Replaces the browser’s default new-tab page with a custom page defined by the researcher.

Good for: Studies that want to present participants with specific content or prompts each time they open a new tab, or studies that need a controlled starting point for browsing sessions.

Limitations: Participants may find a changed new-tab page noticeable or disruptive. Consider disclosure and consent carefully.

Technical reference: See Module Map


What it collects: Nothing.

What it does: Moves each participant through the phases of a study and turns each phase’s settings on automatically. It can also assign participants to conditions at random.

Good for: Multi-phase and experimental studies, such as a baseline period followed by an intervention, or random assignment between a control and treatment group.

Limitations: Phases advance on a timer or when a data upload completes. Changing phase names while a study is running breaks participants’ progress.

Technical reference: See Module Map


Every study is different. The Choosing Modules guide walks through a decision process based on your research questions: what data you need, what level of data reliability you require, and whether your study is observational or experimental.