REX ChatGPT Conversation Data
Schema overview and sample records, prepared by BRIC. Updated September 2026.
Overview
Section titled “Overview”This document describes the ChatGPT conversation data collected by a REX browser extension’s Spider module (rex-spider-chatgpt), including a sample record.
What a record represents
Section titled “What a record represents”A record is one ChatGPT conversation with all of its turns inlined. A conversation is identified by its stable ChatGPT conversation_id; when a participant adds more turns to an existing conversation, the next sync emits a new record under the same identifier reflecting the updated state.
Records are produced by an authenticated API crawl: the spider reads the participant’s conversation list from /backend-api/conversations (including conversations that live inside ChatGPT Projects), then fetches each updated conversation in full. If the API path is unavailable (e.g. the participant is not logged in), the spider falls back to a thinner DOM-scrape record; the fields documented here describe the primary API-crawl record.
At a glance
Section titled “At a glance”Volume: one JSON object per updated ChatGPT conversation. A few to a few dozen records per active day; zero on inactive days.
Granularity: whole conversation, not turn. Turns are inlined as an array; the conversation’s identifier is stable across resyncs. A conversation’s branching and regeneration history is preserved through per-turn parent pointers.
Provenance: every record carries the participant ID, the extension name and version, and a hash of the configuration that produced it in passive-data-metadata, so records remain interpretable after instrumentation changes.
Privacy posture: configurable. The detail_level setting chooses how much of each conversation is collected, from start and end times only up to full text (see Levels of detail). The lookback_days setting bounds how far back the spider will look when enumerating conversations to crawl. Sync frequency is also configurable. The REX Content Processing module can redact content* fields (which carry prompt and response text) before they leave the browser.
Levels of detail
Section titled “Levels of detail”Set detail_level in the spider configuration to collect only what your study needs. Each level is a named set of fields. The levels are not a strict ladder: message_times has no title, and titles has no message times.
| Level | What each conversation record contains |
|---|---|
conversation_times | platform, identifier, started, ended |
message_times | Conversation times, plus one entry per turn with only speaker, when, and identifier. No title and no text. |
titles | Conversation times, plus the conversation title* |
titles_and_message_times | titles plus the message_times turns |
full | Everything shown in the sample record below, including turn text, search results, and citations |
If no level is set, the spider collects full. The older summarize: true setting is the same as conversation_times.
The sample record below is at the full level.
1. Raw record
Section titled “1. Raw record”Below is the JSON for a single two-turn conversation record. All fields present in full production records appear here; nested objects (including each turn’s raw upstream metadata* payload) are shown in full.
{ "name": "rex-conversation", "date": "2026-04-15T14:02:11.000Z", "platform": "chatgpt", "identifier": "abc123de-4567-89ab-cdef-0123456789ab", "title*": "Red Line service check", "started": "2026-04-15T14:02:11Z", "ended": "2026-04-15T14:03:18Z", "turns": [ { "speaker": "user", "when": "2026-04-15T14:02:11Z", "identifier": "aaaaaaa1-1111-2222-3333-444444444444", "content*": "Is the CTA Red Line running normally today, or are there delays?", "parent": "client-created-root", "metadata*": { "id": "aaaaaaa1-1111-2222-3333-444444444444", "message": { "id": "aaaaaaa1-1111-2222-3333-444444444444", "author": { "role": "user", "name": null, "metadata": {} }, "create_time": 1776261731.0, "content": { "content_type": "text", "parts": ["Is the CTA Red Line running normally today, or are there delays?"] }, "status": "finished_successfully", "end_turn": null, "weight": 1.0, "metadata": { "serialization_metadata": { "custom_symbol_offsets": [] } }, "recipient": "all", "channel": null }, "parent": "client-created-root", "children": ["bbbbbbb2-2222-3333-4444-555555555555"] } }, { "speaker": "assistant", "when": "2026-04-15T14:03:18Z", "identifier": "bbbbbbb2-2222-3333-4444-555555555555", "content*": "As of right now, the CTA's service alerts page shows the Red Line running normally. Check transitchicago.com/alerts for the most current updates before you travel.", "parent": "aaaaaaa1-1111-2222-3333-444444444444", "metadata*": { "id": "bbbbbbb2-2222-3333-4444-555555555555", "message": { "id": "bbbbbbb2-2222-3333-4444-555555555555", "author": { "role": "assistant", "name": null, "metadata": {} }, "create_time": 1776261798.0, "content": { "content_type": "text", "parts": ["As of right now, the CTA's service alerts page shows the Red Line running normally. Check transitchicago.com/alerts for the most current updates before you travel."] }, "status": "finished_successfully", "end_turn": true, "weight": 1.0, "metadata": { "model_slug": "gpt-4o", "default_model_slug": "gpt-4o", "parent_id": "aaaaaaa1-1111-2222-3333-444444444444", "search_result_groups": [ { "type": "web", "entries": [ { "title": "CTA Service Alerts", "url": "https://www.transitchicago.com/alerts/", "snippet": "Red Line: normal service. No alerts are in effect.", "ref_id": { "ref_index": 0 } } ] } ] }, "recipient": "all", "channel": null }, "parent": "aaaaaaa1-1111-2222-3333-444444444444", "children": [] }, "search": { "platform": "chatgpt", "query*": "?", "type": "web", "results": [ { "title": "CTA Service Alerts", "url": "https://www.transitchicago.com/alerts/", "preview": "Red Line: normal service. No alerts are in effect.", "index": 0 } ] }, "citations": [ { "title": "CTA Service Alerts", "url": "https://www.transitchicago.com/alerts/", "source": "transitchicago.com" } ] } ], "passive-data-metadata": { "source": "participant-001", "group": "participant-001", "generator": "rex-conversation: Example Study/1.1.0 Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36", "generator-id": "rex-conversation", "timestamp": 1776261731.0, "timezone": "America/Chicago", "configuration-hash": "3f9a1c0e7b2d4a6f8e1b5c9d0a2e4f6b8c1d3e5f7a9b0c2d4e6f8a1b3c5d7e9f", "encrypted_transmission": false }}2. Example parsed view: conversation-level
Section titled “2. Example parsed view: conversation-level”The table below flattens five sample conversation records into an analysis-ready rectangle. Timestamps are shown in the participant’s local timezone (from passive-data-metadata.timezone); the underlying epoch-millisecond values are preserved in the raw data.
| conversation_id | started | ended | platform | turns | models_used | first_prompt_snippet | source |
|---|---|---|---|---|---|---|---|
| abc123de…89ab | 2026-04-15 09:02:11 CT | 2026-04-15 09:03:18 CT | chatgpt | 2 | gpt-4o | Is the CTA Red Line running normally… | participant-001 |
| cd44ef01…12cd | 2026-04-15 09:17:44 CT | 2026-04-15 09:24:02 CT | chatgpt | 8 | gpt-4o, o3-mini | Help me draft a grant cover letter for… | participant-001 |
| e5ff2233…44ef | 2026-04-15 10:41:09 CT | 2026-04-15 10:41:55 CT | chatgpt | 2 | gpt-4o | Translate this paragraph into Spanish: ”…“ | participant-001 |
| f6aa4455…6600 | 2026-04-15 14:08:30 CT | 2026-04-15 14:33:12 CT | chatgpt | 22 | gpt-4o, gpt-4.5 | I’m debugging a Python script that reads… | participant-002 |
| aa77bb88…cc99 | 2026-04-15 16:52:01 CT | 2026-04-15 17:04:47 CT | chatgpt | 6 | o3-mini | Summarize the attached PDF and pull out… | participant-002 |
Columns shown are a curated subset. The full record includes each turn’s content and per-turn metadata (model slug, search results, citations, parent pointers), plus each turn’s raw upstream node; see the field dictionary below.
3. Example parsed view: turn-level
Section titled “3. Example parsed view: turn-level”The same records expanded to one row per turn (showing the first two conversations from the table above).
| conversation_id | turn_identifier | when | speaker | model | content_snippet | has_search | has_citations | source |
|---|---|---|---|---|---|---|---|---|
| abc123de…89ab | aaaaaaa1…4444 | 2026-04-15 09:02:11 CT | user | none | Is the CTA Red Line running normally… | no | no | participant-001 |
| abc123de…89ab | bbbbbbb2…5555 | 2026-04-15 09:03:18 CT | assistant | gpt-4o | As of right now, the CTA’s service alerts… | yes | yes | participant-001 |
| cd44ef01…12cd | 1111aaaa…2222 | 2026-04-15 09:17:44 CT | user | none | Help me draft a grant cover letter for… | no | no | participant-001 |
| cd44ef01…12cd | 2222bbbb…3333 | 2026-04-15 09:18:02 CT | assistant | gpt-4o | Here’s a draft cover letter. I’ve structured… | no | no | participant-001 |
The model column is derived from each assistant turn’s metadata*.message.metadata.model_slug. User turns have no model and are shown as none.
4. Field dictionary
Section titled “4. Field dictionary”Reference for the fields most commonly used in analysis.
Conversation-level fields
Section titled “Conversation-level fields”| Field | Type | Meaning |
|---|---|---|
name | string | Always rex-conversation. Identifies the record type for downstream routing. |
date | ISO timestamp | Conversation start time. Mirrors started and the passive-data-metadata.timestamp. |
platform | string | Always chatgpt for records produced by this spider. |
identifier | string | ChatGPT’s stable conversation_id. Resyncs of the same conversation share this value. |
title* | string | The conversation title as ChatGPT shows it in the sidebar. ChatGPT generates titles from the conversation, so they can contain personal details; the asterisk marks the field for redaction. Present at the titles, titles_and_message_times, and full levels. |
started | ISO timestamp | The earliest turn’s create time. |
ended | ISO timestamp | The latest turn’s create time. |
turns | array | Ordered turns in the conversation. Present at the message_times, titles_and_message_times, and full levels. See turn-level fields below. |
metadata | object | ChatGPT’s raw conversation payload, including the full mapping tree. Only present when the spider runs in debug mode, which is for development, not studies. |
passive-data-metadata | object | Capture context: source (participant ID), group, generator (record type, extension name and version, and browser user agent), generator-id, timestamp (seconds since 1970), timezone, configuration-hash, and encrypted_transmission. |
Turn-level fields
Section titled “Turn-level fields”| Field | Type | Meaning |
|---|---|---|
speaker | string | user, assistant, system, or tool, passed through from the upstream author.role. |
when | ISO timestamp | The turn’s create time. |
identifier | string | ChatGPT’s message ID for this turn. |
content* | string | Turn text. For multi-part messages, parts are joined with newlines. Subject to redaction via REX Content Processing. |
parent | string | Identifier of the parent turn, or client-created-root for the synthetic root. Use these pointers to reconstruct the conversation tree (branches and regenerations). |
metadata* | object | Raw upstream node from the conversation’s mapping. Includes message.metadata.model_slug, search_result_groups, message.status, and children. full level only. |
search | object | Present when the assistant’s turn triggered a web search. Carries platform, type, and a results array. |
citations | array | Present when the assistant’s turn cited sources. Each citation has title, url, and source. |
Notes on branches and parent pointers
Section titled “Notes on branches and parent pointers”ChatGPT stores a conversation as a tree, not a flat list, so that branches (regenerations, edits to earlier turns) can coexist. The turns array is a pre-order traversal of that tree, so turns from abandoned branches appear alongside the “active” branch, and parent pointers are needed to distinguish them. For most analyses, start from the latest turn (by when) that has no children in its metadata*, then walk parent pointers back to the root. That gives you the branch the participant saw last.
Notes on the * suffix
Section titled “Notes on the * suffix”Fields whose names end in * (title*, content*, metadata*, query*) carry content that may include participant prompts, model responses, or upstream payloads that embed them. The REX Content Processing module can be configured to redact these fields before transmission. Fields without the * suffix (identifiers, timestamps, model slugs, URLs in citations) are not subject to that redaction pass.