Status: Infrastructure built (2026-10-03, task #705) — backend pipeline, parsers (verified against real exports from Anthropic, Google, xAI), review/commit flow, and the SPA wizard at System → Import are live. The AI analysis stage (summaries + proposed items) is the remaining piece; the ChatGPT sample export is still gathering. We ship two open standards as import AND export, out of the box: the Konshus AI Memory Portability Standard v0.1 (atoms + artifacts — your distilled memory) and the Data Transfer Initiative's AI Conversation History schema (schemas.pub/schemas/24 — message-level conversation interchange). Once your memory lives in open formats, provider schema churn stops being your problem: any v0.1- or DTI-speaking tool can take it from here, and we can ingest theirs. Everything below reflects the BUILT reality, not the original plan — the original format section was substantially wrong about what providers actually ship.

Data Import Pipeline

Your conversations aren't just chat logs -- they're the record of a relationship. When you leave a provider, you shouldn't have to start that relationship from scratch. Export your history, upload it to Atamaia, and your AI picks up right where you left off -- with any model you choose.


The Opportunity

700K+ users are leaving ChatGPT. Many will jump to Claude, but Anthropic's pricing will cause churn too. These users have months or years of conversation history -- preferences learned, projects discussed, communication styles established. Every platform they move to makes them start from zero.

Atamaia is the permanent home. The killer onboarding: export your chats from any provider, upload them, and we analyze them to build your AI identity and seed your memory system. Five minutes from "I just quit ChatGPT" to "my AI already knows me."

What the providers actually give you (verified 2026-10-03 against real exports)

Every provider makes leaving harder than staying. The hoops are worth recording, because the pipeline has to be built around the real shapes, not the polite ones:

Provider What the export actually is The hoops
ChatGPT (OpenAI) ZIP with conversations.json (tree structure) + chat.html The request itself is a gauntlet: sign-in, emailed code, then a separate confirm screen before the export even queues — then the link arrives by email, often days later (some users report it never arrives). The export omits Memory entries and DALL-E images entirely, and cannot be requested from the mobile app
Claude (Anthropic) A manifest (anthropic.json) listing FIVE category ZIPs: light_metadata, projects, memories, frames, conversations. Conversations are flat chat_messages; projects ship their docs with full content; claude.ai's own memory blob ships in the memories part Links expire in 24 hours and may work only once (the manifest says so itself); each part downloads separately
Gemini (Google Takeout) An activity log, not chats. Takeout/My Activity/Gemini Apps/My Activity.json — each entry is one turn: the user prompt survives only as the (possibly truncated) entry title ("Prompted …"), the response only as HTML-escaped HTML (safeHtmlItem), and the conversation exists only as a gemini.google.com URL The "Gemini" checkbox in Takeout exports custom Gems only, not chats (you have to select My Activity → Gemini Apps instead); prompt text is truncated by Google — data you cannot get out in full
Grok (xAI) An internal service dump: prod-grok-backend.json (Mongo extended JSON, ~43MB) with conversations (response tree), projects, tasks; attachments are anonymous uuid blobs under prod-mc-asset-server/ with no filenames; everything under a ttl/30d/ path No friendly schema at all — you get their internal database shape, opaque asset blobs, and an export that itself expires

Rich's own first Gemini activity entry is the thesis in one line: "I need to bulk download my gemini chats - not one by one, and takeout doesn't work with chats" — captured inside the very export that doesn't work.

Media: the blind spot everywhere

The text is hard enough to get out; media is worse. ChatGPT's export omits generated images entirely. Gemini's Takeout dumps images and documents as flat files with hash-suffixed names and no link back to the conversation they belonged to. Grok's export stores attachments as anonymous uuid blobs with no filenames or types. Claude references attached files by uuid and ships project documents separately. Anyone importing "the whole relationship" is really importing its text — every provider makes the pictures harder to keep than the words.

On the platform side, the consumer memory tools are text-first too: mem0 and Letta ingest text (mem0 through an add() extraction API, Letta through per-passage archival inserts) and neither documents media import. Konshus is the exception worth studying — its Shoebox feature runs photos and handwritten pages through vision models and files the results as sourced, reviewable atoms. That is the right instinct, and it is where our own reserved Media and Artifact item types are pointed: media needs file storage and a vision extraction pass before it can join the review flow, and both are on the roadmap rather than in it today.

How we keep up with the revolving door

Provider schemas drift; new providers appear. The strategy is parsers as code, not a transform configuration language: one parser per provider behind one interface, structural detection that fails loudly on unknown shapes ("matched no supported format" beats silently mis-parsing), fixture tests per shape so a schema change is a new fixture next to the old regression, and a format/version stamp on every import row so we always know which parser revision produced which staged data. The open Konshus format is the long-game answer to the churn itself.


The User Journey

Step 1: Export Your Data

From ChatGPT (OpenAI)

  1. Settings > Data Controls > Export Data > confirm via email
  2. Download the ZIP (conversations.json is the gold)

From Claude (Anthropic)

  1. Settings > Privacy > Export Data
  2. The manifest (anthropic.json) lists the category parts — download ALL of them promptly; links expire in 24h and may work only once
  3. conversations-000.zip is the chat history; projects-000.zip carries projects + docs; memories-000.zip carries claude.ai's memory

From Google Gemini

  1. takeout.google.com > Deselect all > My Activity > All activity data included > Gemini Apps only
  2. Do NOT rely on the "Gemini" checkbox — it exports Gems and extensions, not chats
  3. Download when ready (hours for large accounts)

From Grok (xAI)

  1. Settings → download your data
  2. The archive is an internal dump — we parse prod-grok-backend.json directly

Step 2: Upload (System → Import)

Drag the ZIP onto the wizard, pick the identity that will receive memories and (optionally) the project that will receive facts. Auto-detection is structural — the pipeline looks at the archive's shape, not the file name.

Step 3: Parse and Scrub (automatic)

The pipeline parses and normalizes entirely on local infrastructure — no cloud APIs:

  • Detects the provider format from archive structure
  • Normalizes all formats to one internal conversation shape
  • Scrubs for secrets and obvious PII before anything reaches review: API keys (OpenAI/Anthropic/GitHub/Slack/Google/AWS patterns), JWTs, private-key blocks, credential assignments, email addresses. Findings are attached to the flagged content as MASKED excerpts — the flag never contains the secret itself, but you can find the spot and decide
  • Stages everything for review; nothing touches memories or facts yet

Step 4: Review — the point of the whole feature

Nothing is committed without an explicit decision, per conversation and per item:

  • Conversations: approve (becomes a Conversation memory on the target identity), discard, or read the transcript inline first
  • Scrub warnings show which conversations contain likely secrets — "Approve all without flags" in one click
  • Extracted items (facts, proposed memories, preferences, identity traits): approve, edit, or discard each one
  • Provider warnings are honest: Gemini conversations say upfront that Google's export truncated your prompts; Grok conversations say attachments aren't carried yet

Step 5: Commit

Approved conversations become memories (summary preferred over raw transcript; transcripts are truncated at 4,000 characters with an explicit marker rather than embedded verbatim — a full transcript embedded into the vector store is exactly the contamination this pipeline exists to prevent). Approved facts/preferences land in the chosen project's fact store. Memory provenance is Reported — relayed from an external source, not asserted — so downstream weighting can tell imported content from native memory.


Parsed formats (verified against real exports)

ChatGPT / OpenAI — conversations.json

Tree-structured: mapping links message nodes by parent/children UUIDs; branching happens when the user edits or regenerates. The pipeline reconstructs the live branch by walking from current_node back to root and reversing — abandoned regenerations are not history the user kept. Fields: title, create_time/update_time (Unix epoch floats), per-message author.role (system/user/assistant/tool), content.content_type + content.parts (string array), metadata.model_slug.

Claude / Anthropic — conversations.json (conversations part)

Flat: uuid, name, summary, created_at/updated_at (ISO 8601), chat_messages[] with sender (human/assistant), text, content[] blocks. Content blocks include text, thinking (the model's private reasoning — skipped), tool_use/tool_result (skipped in v1). Attachments carry inline extracted_content; files are referenced by uuid only.

Claude projects part — projects/<uuid>.json

Per project: name, description, prompt_template, docs[] with full content (uuid, filename, content, created_at). A direct map onto Atamaia projects + docs. (Parsed in a following increment — the triage targets exist, the parser is next.)

Claude memories part — memories/<uuid>.json

claude.ai's own accumulated memory as one conversations_memory blob. Maps to candidate IdentityTrait/semantic items for review. (Following increment.)

Google Gemini — Takeout/My Activity/Gemini Apps/My Activity.json

Activity log, not conversations. 161 entries grouped into 21 conversations by their gemini.google.com URL slug in details[0].name. Per entry: title = "Prompted " + prompt text (Google may truncate), time, safeHtmlItem[].html = the response as HTML-escaped HTML (tags stripped, entities decoded at parse). "Cleared previous feedback"-style non-turn entries are skipped. Consecutive same-speaker turns are merged. Prompt truncation is a data-loss property of Google's export and is warned about in the UI — the pipeline will not present a rebuilt transcript as if it were complete.

Grok — prod-grok-backend.json

Internal dump, Mongo extended JSON ({"$date": {"$numberLong": "millis"}} timestamps). conversations[] of {conversation, responses[]} where responses form a tree (parent_response_id, children, leaf_response_id) — reconstructed the same way as ChatGPT's live branch. sender human/otherwise; model on responses; conversation title, create_time, modify_time. Projects exist with custom_personality + conversation_starters. Attachments are uuid keys into prod-mc-asset-server/<uuid>/content blobs — media handling is a following increment.

Other providers

Provider Export exists? Parser
Microsoft Copilot Privacy dashboard JSON (not yet sampled) pending sample
Perplexity No official export monitor

Built architecture (2026-10-03)

API (permission: ImportData, the Data Portability sibling of ExportData)

POST   /api/import/upload                          multipart: file + identityId? + projectId?
GET    /api/import                                 paged list
GET    /api/import/{id}                            status (poll while Parsing/Committing)
GET    /api/import/{id}/review                     staged conversations + items + scrub flags
GET    /api/import/{id}/conversations/{cid}        full transcript of one staged conversation
PATCH  /api/import/{id}/review                     triage decisions (bulk)
POST   /api/import/{id}/commit                     begin committing (detached; retarget identity/project allowed)
DELETE /api/import/{id}                            discard (soft delete)

The identity-access guard (#401) runs on every identityId in play: ImportData answers "may import", never "import WHOSE" — committing into an identity's memory is the write-direction equivalent of exporting it.

Pipeline

Upload stores the archive under the import scratch root (default /tmp/atamaia-imports, Import:ScratchRoot) and processing runs detached (the EmbeddingBackfillService shape): parse → detect → normalize → scrub → stage → ReadyForReview. The archive is deleted once parsed — processed and discarded, win or lose. Commit runs detached too (embeddings generate inline per memory) and stamps provenance (CommittedMemoryId / CommittedFactId) on each row it wrote.

Schema

imports (status/provider/counts/scrub totals/error), import_conversations (normalized transcript as JSONB, scrub flags as JSONB, approved, summary, committed_memory_id), import_items (staged extracted data: type/key/value JSONB/confidence/decision/edited value/commit provenance). Enums as seeded lookups (D5): import_statuses, import_source_providers, import_item_types, import_item_decisions.

SPA

System → Import: upload (drag-drop + identity/project pickers) → polling while Parsing → review (scrub badges, inline transcripts, approve-all-clean, extracted-item triage) → commit → done. The review step is the product: nothing moves to permanent storage without the user's per-item decision.


The analysis stage (next increment)

The AI pass that Aether and Rich sketched (2026-10-03, chat session 81): per-conversation AI summaries and candidate item extraction — sanitize → AI synthesize → human triage → route to storage tier. Precedents already studied there: Microsoft Recall (sensitive filtering + opt-outs), MemGPT/Letta (an agentic summarizer proposes memory edits rather than dumping raw logs), Readwise-style tag-and-discard ingestion, DLP tooling (regex + entropy + small NER) for the boundary pass — the scrubber is the built v1 of that last one.

Design constraints already decided:

  • Analysis runs on the local analysis model (resolved by name, never a hardcoded id) — zero cloud calls, zero token cost
  • Chunked: ~20 conversations per call, parallel batches, aggregate + dedupe + confidence-rank, one final refinement pass
  • Everything the analysis proposes lands in import_items as PENDING — it proposes, the human decides

Privacy & security

  • All processing local: no third-party API calls during import or analysis
  • Scrub flags surface secrets BEFORE commit; the user decides per item (the scrubber flags, never deletes)
  • Nothing commits without explicit approval, per conversation and per item
  • The uploaded archive is deleted after parse (success or failure)
  • Raw staged transcripts are soft-deleted at commit (D15) — they leave review, but the rows persist for audit. Doc 80's original promise said raw conversations would not be stored permanently; making that a hard delete trades undo-ability for the privacy promise, and that trade belongs to a deliberate decision, not a silent one
  • Multi-tenant isolation throughout; identities can only be import targets if the caller may act for them