Skip to content

REX ChatGPT Conversation Data

Schema overview and sample records, prepared by BRIC. Updated September 2026.

This document describes the ChatGPT conversation data collected by a REX browser extension’s Spider module (rex-spider-chatgpt), including a sample record.

A record is one ChatGPT conversation with all of its turns inlined. A conversation is identified by its stable ChatGPT conversation_id; when a participant adds more turns to an existing conversation, the next sync emits a new record under the same identifier reflecting the updated state.

Records are produced by an authenticated API crawl: the spider reads the participant’s conversation list from /backend-api/conversations (including conversations that live inside ChatGPT Projects), then fetches each updated conversation in full. If the API path is unavailable (e.g. the participant is not logged in), the spider falls back to a thinner DOM-scrape record; the fields documented here describe the primary API-crawl record.

Volume: one JSON object per updated ChatGPT conversation. A few to a few dozen records per active day; zero on inactive days.

Granularity: whole conversation, not turn. Turns are inlined as an array; the conversation’s identifier is stable across resyncs. A conversation’s branching and regeneration history is preserved through per-turn parent pointers.

Provenance: every record carries the participant ID, the extension name and version, and a hash of the configuration that produced it in passive-data-metadata, so records remain interpretable after instrumentation changes.

Privacy posture: configurable. The detail_level setting chooses how much of each conversation is collected, from start and end times only up to full text (see Levels of detail). The lookback_days setting bounds how far back the spider will look when enumerating conversations to crawl. Sync frequency is also configurable. The REX Content Processing module can redact content* fields (which carry prompt and response text) before they leave the browser.

Set detail_level in the spider configuration to collect only what your study needs. Each level is a named set of fields. The levels are not a strict ladder: message_times has no title, and titles has no message times.

LevelWhat each conversation record contains
conversation_timesplatform, identifier, started, ended
message_timesConversation times, plus one entry per turn with only speaker, when, and identifier. No title and no text.
titlesConversation times, plus the conversation title*
titles_and_message_timestitles plus the message_times turns
fullEverything shown in the sample record below, including turn text, search results, and citations

If no level is set, the spider collects full. The older summarize: true setting is the same as conversation_times.

The sample record below is at the full level.


Below is the JSON for a single two-turn conversation record. All fields present in full production records appear here; nested objects (including each turn’s raw upstream metadata* payload) are shown in full.

{
"name": "rex-conversation",
"date": "2026-04-15T14:02:11.000Z",
"platform": "chatgpt",
"identifier": "abc123de-4567-89ab-cdef-0123456789ab",
"title*": "Red Line service check",
"started": "2026-04-15T14:02:11Z",
"ended": "2026-04-15T14:03:18Z",
"turns": [
{
"speaker": "user",
"when": "2026-04-15T14:02:11Z",
"identifier": "aaaaaaa1-1111-2222-3333-444444444444",
"content*": "Is the CTA Red Line running normally today, or are there delays?",
"parent": "client-created-root",
"metadata*": {
"id": "aaaaaaa1-1111-2222-3333-444444444444",
"message": {
"id": "aaaaaaa1-1111-2222-3333-444444444444",
"author": { "role": "user", "name": null, "metadata": {} },
"create_time": 1776261731.0,
"content": {
"content_type": "text",
"parts": ["Is the CTA Red Line running normally today, or are there delays?"]
},
"status": "finished_successfully",
"end_turn": null,
"weight": 1.0,
"metadata": { "serialization_metadata": { "custom_symbol_offsets": [] } },
"recipient": "all",
"channel": null
},
"parent": "client-created-root",
"children": ["bbbbbbb2-2222-3333-4444-555555555555"]
}
},
{
"speaker": "assistant",
"when": "2026-04-15T14:03:18Z",
"identifier": "bbbbbbb2-2222-3333-4444-555555555555",
"content*": "As of right now, the CTA's service alerts page shows the Red Line running normally. Check transitchicago.com/alerts for the most current updates before you travel.",
"parent": "aaaaaaa1-1111-2222-3333-444444444444",
"metadata*": {
"id": "bbbbbbb2-2222-3333-4444-555555555555",
"message": {
"id": "bbbbbbb2-2222-3333-4444-555555555555",
"author": { "role": "assistant", "name": null, "metadata": {} },
"create_time": 1776261798.0,
"content": {
"content_type": "text",
"parts": ["As of right now, the CTA's service alerts page shows the Red Line running normally. Check transitchicago.com/alerts for the most current updates before you travel."]
},
"status": "finished_successfully",
"end_turn": true,
"weight": 1.0,
"metadata": {
"model_slug": "gpt-4o",
"default_model_slug": "gpt-4o",
"parent_id": "aaaaaaa1-1111-2222-3333-444444444444",
"search_result_groups": [
{
"type": "web",
"entries": [
{
"title": "CTA Service Alerts",
"url": "https://www.transitchicago.com/alerts/",
"snippet": "Red Line: normal service. No alerts are in effect.",
"ref_id": { "ref_index": 0 }
}
]
}
]
},
"recipient": "all",
"channel": null
},
"parent": "aaaaaaa1-1111-2222-3333-444444444444",
"children": []
},
"search": {
"platform": "chatgpt",
"query*": "?",
"type": "web",
"results": [
{
"title": "CTA Service Alerts",
"url": "https://www.transitchicago.com/alerts/",
"preview": "Red Line: normal service. No alerts are in effect.",
"index": 0
}
]
},
"citations": [
{
"title": "CTA Service Alerts",
"url": "https://www.transitchicago.com/alerts/",
"source": "transitchicago.com"
}
]
}
],
"passive-data-metadata": {
"source": "participant-001",
"group": "participant-001",
"generator": "rex-conversation: Example Study/1.1.0 Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
"generator-id": "rex-conversation",
"timestamp": 1776261731.0,
"timezone": "America/Chicago",
"configuration-hash": "3f9a1c0e7b2d4a6f8e1b5c9d0a2e4f6b8c1d3e5f7a9b0c2d4e6f8a1b3c5d7e9f",
"encrypted_transmission": false
}
}

2. Example parsed view: conversation-level

Section titled “2. Example parsed view: conversation-level”

The table below flattens five sample conversation records into an analysis-ready rectangle. Timestamps are shown in the participant’s local timezone (from passive-data-metadata.timezone); the underlying epoch-millisecond values are preserved in the raw data.

conversation_idstartedendedplatformturnsmodels_usedfirst_prompt_snippetsource
abc123de…89ab2026-04-15 09:02:11 CT2026-04-15 09:03:18 CTchatgpt2gpt-4oIs the CTA Red Line running normally…participant-001
cd44ef01…12cd2026-04-15 09:17:44 CT2026-04-15 09:24:02 CTchatgpt8gpt-4o, o3-miniHelp me draft a grant cover letter for…participant-001
e5ff2233…44ef2026-04-15 10:41:09 CT2026-04-15 10:41:55 CTchatgpt2gpt-4oTranslate this paragraph into Spanish: ”…“participant-001
f6aa4455…66002026-04-15 14:08:30 CT2026-04-15 14:33:12 CTchatgpt22gpt-4o, gpt-4.5I’m debugging a Python script that reads…participant-002
aa77bb88…cc992026-04-15 16:52:01 CT2026-04-15 17:04:47 CTchatgpt6o3-miniSummarize the attached PDF and pull out…participant-002

Columns shown are a curated subset. The full record includes each turn’s content and per-turn metadata (model slug, search results, citations, parent pointers), plus each turn’s raw upstream node; see the field dictionary below.


The same records expanded to one row per turn (showing the first two conversations from the table above).

conversation_idturn_identifierwhenspeakermodelcontent_snippethas_searchhas_citationssource
abc123de…89abaaaaaaa1…44442026-04-15 09:02:11 CTusernoneIs the CTA Red Line running normally…nonoparticipant-001
abc123de…89abbbbbbbb2…55552026-04-15 09:03:18 CTassistantgpt-4oAs of right now, the CTA’s service alerts…yesyesparticipant-001
cd44ef01…12cd1111aaaa…22222026-04-15 09:17:44 CTusernoneHelp me draft a grant cover letter for…nonoparticipant-001
cd44ef01…12cd2222bbbb…33332026-04-15 09:18:02 CTassistantgpt-4oHere’s a draft cover letter. I’ve structured…nonoparticipant-001

The model column is derived from each assistant turn’s metadata*.message.metadata.model_slug. User turns have no model and are shown as none.


Reference for the fields most commonly used in analysis.

FieldTypeMeaning
namestringAlways rex-conversation. Identifies the record type for downstream routing.
dateISO timestampConversation start time. Mirrors started and the passive-data-metadata.timestamp.
platformstringAlways chatgpt for records produced by this spider.
identifierstringChatGPT’s stable conversation_id. Resyncs of the same conversation share this value.
title*stringThe conversation title as ChatGPT shows it in the sidebar. ChatGPT generates titles from the conversation, so they can contain personal details; the asterisk marks the field for redaction. Present at the titles, titles_and_message_times, and full levels.
startedISO timestampThe earliest turn’s create time.
endedISO timestampThe latest turn’s create time.
turnsarrayOrdered turns in the conversation. Present at the message_times, titles_and_message_times, and full levels. See turn-level fields below.
metadataobjectChatGPT’s raw conversation payload, including the full mapping tree. Only present when the spider runs in debug mode, which is for development, not studies.
passive-data-metadataobjectCapture context: source (participant ID), group, generator (record type, extension name and version, and browser user agent), generator-id, timestamp (seconds since 1970), timezone, configuration-hash, and encrypted_transmission.
FieldTypeMeaning
speakerstringuser, assistant, system, or tool, passed through from the upstream author.role.
whenISO timestampThe turn’s create time.
identifierstringChatGPT’s message ID for this turn.
content*stringTurn text. For multi-part messages, parts are joined with newlines. Subject to redaction via REX Content Processing.
parentstringIdentifier of the parent turn, or client-created-root for the synthetic root. Use these pointers to reconstruct the conversation tree (branches and regenerations).
metadata*objectRaw upstream node from the conversation’s mapping. Includes message.metadata.model_slug, search_result_groups, message.status, and children. full level only.
searchobjectPresent when the assistant’s turn triggered a web search. Carries platform, type, and a results array.
citationsarrayPresent when the assistant’s turn cited sources. Each citation has title, url, and source.

ChatGPT stores a conversation as a tree, not a flat list, so that branches (regenerations, edits to earlier turns) can coexist. The turns array is a pre-order traversal of that tree, so turns from abandoned branches appear alongside the “active” branch, and parent pointers are needed to distinguish them. For most analyses, start from the latest turn (by when) that has no children in its metadata*, then walk parent pointers back to the root. That gives you the branch the participant saw last.

Fields whose names end in * (title*, content*, metadata*, query*) carry content that may include participant prompts, model responses, or upstream payloads that embed them. The REX Content Processing module can be configured to redact these fields before transmission. Fields without the * suffix (identifiers, timestamps, model slugs, URLs in citations) are not subject to that redaction pass.