Instagram Comments and Commenter Profiles for Footwear Audience Analysis
A footwear research project needed clean Instagram source data for downstream analysis of audience overlap, engagement patterns, and internal influence. Orzaen delivered a full brand post index, then structured top-level comments and unique public commenter profiles from a client-selected 38-post subset in separate joinable datasets.
Case visual
Add gallery images to show system screenshots, dashboards, workflow outputs, or delivery samples.
Approved Post Subset
Linked Data Layers
Comment Scope
Analysis Boundary
The client was not asking for a generic comment export. The intended research depended on selecting the right posts, preserving raw engagement records, and linking those records to public commenter profiles without mixing extraction with interpretation.
Orzaen created the full post index first, then processed the client's 38-post subset into separate comments and commenter-profile files. This staged structure produced a frozen pilot schema that could be reused across additional footwear accounts after validation.
Behind the Scenes
How the system moved from problem to controlled execution.
Problem
The client planned a long-running comparison of footwear-brand audiences, but first needed a pilot that proved the underlying Instagram data could be collected cleanly and joined across extraction stages. The project required a complete post index before the client selected high-signal posts. Comments and commenter profiles needed separate schemas without losing their relationships to the source posts. Repeated commenters had to be represented once in the profile dataset. The client would perform all scoring and interpretation, so Orzaen had to preserve public source data without enrichment, filtering, or assumptions. The approved deep-extraction subset contained exactly 38 posts.
System Built
Orzaen separated the pilot into a full post-index stage and a selected-post deep-extraction stage. The post index preserved public post identifiers, URLs, dates, captions, and engagement counts so the client could choose the relevant subset. For the 38 approved posts, Orzaen collected the visible top-level comments into a raw comment dataset and built a separate unique commenter-profile dataset containing available public identity, biography, activity, link, and account-status fields. Stable post identifiers and usernames kept the datasets joinable while analysis remained entirely client-led.
What Changed
A full post index gave the client control over post selection. Exactly 38 approved posts were processed in the pilot deep extraction. Comments and unique commenter profiles were delivered as separate, joinable datasets. The schema was structured for reuse across later footwear-brand accounts. All scoring, overlap analysis, and interpretation remained with the client.
Before / After
What changed after the system was rebuilt.
Post selection
Before
No structured view of the brand's post history
After
Full post index available before deep extraction
Comment data
Before
Engagement visible only within individual posts
After
Top-level comments structured across 38 approved posts
Commenter identities
Before
The same person could appear in several comment threads
After
One public profile row per unique commenter
Dataset relationships
Before
Posts, comments, and profiles not available as joinable tables
After
Stable identifiers connect all three data layers
Delivery Scope
What was included in the system delivery.
Full Instagram post index for the pilot brand
Selected 38-post extraction manifest
UTF-8 raw top-level comments CSV
UTF-8 unique commenter-profiles CSV
Stable join fields across posts, comments, and profiles
Reusable frozen schema for later brand-account runs
Controls
Checks built in to keep the workflow reliable.
The post-index and deep-extraction stages remain separate
Only the 38 client-approved posts are included in the deep extraction
Top-level comments are collected; replies are excluded by confirmed scope
Unique commenter profiles are deduplicated within the approved brand dataset
No enrichment, filtering, scoring, or assumptions are added
Public source values are retained in the frozen schema
Tools & Stack
Tools used to build, connect, and deliver this system
Similar System
Want similar results?
Share the manual process, messy data flow, or system gap you want to fix. We will help you understand what can be rebuilt into a controlled operating system.
