Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen
Back to Results

Instagram Comments and Commenter Profiles for Footwear Audience Analysis

A footwear research project needed clean Instagram source data for downstream analysis of audience overlap, engagement patterns, and internal influence. Orzaen delivered a full brand post index, then structured top-level comments and unique public commenter profiles from a client-selected 38-post subset in separate joinable datasets.

Delivered Pilot delivered in staged checkpoints•Footwear Audience Research Client (Confidential)•Pilot delivered in staged checkpoints

Case visual

Add gallery images to show system screenshots, dashboards, workflow outputs, or delivery samples.

IndustryFootwear & Fashion· Consumer Research
38 posts

Approved Post Subset

Posts + comments + profiles

Linked Data Layers

Top-level comments

Comment Scope

Client-led

Analysis Boundary

The client was not asking for a generic comment export. The intended research depended on selecting the right posts, preserving raw engagement records, and linking those records to public commenter profiles without mixing extraction with interpretation.

Orzaen created the full post index first, then processed the client's 38-post subset into separate comments and commenter-profile files. This staged structure produced a frozen pilot schema that could be reused across additional footwear accounts after validation.

Behind the Scenes

How the system moved from problem to controlled execution.

01

Problem

The client planned a long-running comparison of footwear-brand audiences, but first needed a pilot that proved the underlying Instagram data could be collected cleanly and joined across extraction stages. The project required a complete post index before the client selected high-signal posts. Comments and commenter profiles needed separate schemas without losing their relationships to the source posts. Repeated commenters had to be represented once in the profile dataset. The client would perform all scoring and interpretation, so Orzaen had to preserve public source data without enrichment, filtering, or assumptions. The approved deep-extraction subset contained exactly 38 posts.

02

System Built

Orzaen separated the pilot into a full post-index stage and a selected-post deep-extraction stage. The post index preserved public post identifiers, URLs, dates, captions, and engagement counts so the client could choose the relevant subset. For the 38 approved posts, Orzaen collected the visible top-level comments into a raw comment dataset and built a separate unique commenter-profile dataset containing available public identity, biography, activity, link, and account-status fields. Stable post identifiers and usernames kept the datasets joinable while analysis remained entirely client-led.

03

What Changed

A full post index gave the client control over post selection. Exactly 38 approved posts were processed in the pilot deep extraction. Comments and unique commenter profiles were delivered as separate, joinable datasets. The schema was structured for reuse across later footwear-brand accounts. All scoring, overlap analysis, and interpretation remained with the client.

Before / After

What changed after the system was rebuilt.

01

Post selection

Before

No structured view of the brand's post history

After

Full post index available before deep extraction

02

Comment data

Before

Engagement visible only within individual posts

After

Top-level comments structured across 38 approved posts

03

Commenter identities

Before

The same person could appear in several comment threads

After

One public profile row per unique commenter

04

Dataset relationships

Before

Posts, comments, and profiles not available as joinable tables

After

Stable identifiers connect all three data layers

Delivery Scope

What was included in the system delivery.

Full Instagram post index for the pilot brand

Selected 38-post extraction manifest

UTF-8 raw top-level comments CSV

UTF-8 unique commenter-profiles CSV

Stable join fields across posts, comments, and profiles

Reusable frozen schema for later brand-account runs

Controls

Checks built in to keep the workflow reliable.

The post-index and deep-extraction stages remain separate

Only the 38 client-approved posts are included in the deep extraction

Top-level comments are collected; replies are excluded by confirmed scope

Unique commenter profiles are deduplicated within the approved brand dataset

No enrichment, filtering, scoring, or assumptions are added

Public source values are retained in the frozen schema

Tools & Stack

Tools used to build, connect, and deliver this system

PPython
RARequests and browser-based extraction
SASession and rate-limit controls
PAPost and comment parsing
CDCommenter-profile deduplication
UCUTF-8 CSV delivery

Similar System

Want similar results?

Share the manual process, messy data flow, or system gap you want to fix. We will help you understand what can be rebuilt into a controlled operating system.

Achieve Similar Results