Short answer
A reliable Instagram public-profile schema separates identity, source membership, raw profile observations, parsed biography elements, public contact fields, content relationships, and QA metadata. Do not place every field into one undocumented “lead” row: usernames change, biographies contain free text, one profile can appear under several source accounts, and contact values can be missing or ambiguous.
Four questions every value should answer
The schema should answer four questions for every value:
- Which Instagram identity does it describe?
- Which source account or post placed the identity in scope?
- Was the value published directly, parsed, or normalized?
- When was it collected?
The four data grains
Instagram audience projects commonly contain four different grains.
Source account
One row per approved account whose followers, following, posts, or comments are in scope.
Profile
One row per resolved Instagram identity.
Membership or appearance
One row per relationship between a profile and a source account, post, or interaction type.
Content or interaction
One row per post, comment, reply, or other approved public interaction.
Mixing these grains causes duplication. If a profile wrote ten comments, a flat file may repeat its biography, website, and follower count ten times. If that profile also follows three source accounts, the same fields may be repeated again.
Recommended relational model
source_accounts
source_account_id
source_username
source_profile_url
source_category
population_type
displayed_count
requested_limit
collection_started_at
collection_completed_at
collection_status
status_noteprofiles
master_profile_id
instagram_user_id
username_raw
username_current
display_name_raw
profile_url
is_private
is_verified
profile_status
profile_collected_atprofile_observations
master_profile_id
observed_at
biography_raw
followers_count
following_count
posts_count
business_category_raw
external_url_raw
public_email_raw
public_phone_rawaudience_memberships
source_account_id
master_profile_id
relationship_type
observed_atPossible relationship_type values include:
follower
following
commenter
replier
liker
mentioned_profile
tagged_profile
client_providedprofile_links
master_profile_id
link_position
link_url_raw
link_url_normalized
link_domain
link_type
observed_atposts
post_id
source_account_id
post_url
post_type
caption_raw
published_at
public_like_count
public_comment_count
collected_atcomments
comment_id
post_id
commenter_profile_id
parent_comment_id
comment_text_raw
commented_at
public_like_count
collected_atThis model supports follower exports, overlap analysis, comment research, and repeated profile snapshots without overwriting history.
A practical flat-file schema
Not every project needs a database. A well-defined Excel or CSV delivery can still be reliable.
Recommended master fields:
| Field | Meaning |
|---|---|
master_profile_id | Client-side stable identifier |
instagram_user_id | Platform identity when available |
username | Username observed during enrichment |
display_name_raw | Display name exactly as published |
biography_raw | Full biography before parsing |
bio_hashtags | Hashtags parsed from biography |
bio_mentions | Mentions parsed from biography |
followers_count | Public count observed during collection |
following_count | Public count observed during collection |
posts_count | Public count observed during collection |
is_private | Profile privacy status observed |
is_verified | Verification status observed |
business_category_raw | Category as published |
public_email_raw | Email visibly published when available |
public_phone_raw | Phone visibly published when available |
bio_links | Public links retained in order |
source_accounts | Accounts that placed the profile in scope |
source_count | Number of unique source memberships |
profile_url | Public profile URL |
profile_collected_at | Time of the profile observation |
qa_status | Delivery review status |
Preserve raw and parsed biography fields
An Instagram biography is free text. It may contain:
- names and roles;
- emoji;
- hashtags;
- mentions;
- location phrases;
- business claims;
- links typed as text;
- line breaks;
- several languages.
Always retain biography_raw before creating parsed fields.
Example:
biography_raw: "Independent studio | Chicago + NYC | @projectname | #slowfashion"
bio_mentions: "projectname"
bio_hashtags: "slowfashion"
location_text_raw: "Chicago + NYC"Do not silently turn Chicago + NYC into a verified city and state. It may describe markets served, previous residence, travel, or something else. If location is important, define a conservative parsing rule and retain the raw text.
Model multiple bio links properly
A profile may publish:
- one external URL;
- a link-in-bio page containing several destinations;
- a booking link;
- a portfolio;
- an online store;
- a newsletter;
- another social profile.
For a spreadsheet, fixed columns can be practical:
bio_link_1
bio_link_2
bio_link_3
bio_link_4
bio_link_5
bio_link_6This matches Orzaen's six-account entertainment project, which used a 17-field public profile schema with up to six links per profile. See the multi-account entertainment audience case study.
For a database, use a separate profile_links table. Preserve link order and raw URL before normalization.
Do not overwrite raw values with normalized values
Examples:
username_raw -> username_normalized
display_name_raw -> display_name_search
public_phone_raw -> public_phone_e164
link_url_raw -> link_url_normalized
business_category_raw -> business_category_mappedNormalization supports search and comparison, but the raw value is the source observation. If a mapping rule changes later, the record can be reprocessed without revisiting the profile.
Model missing values explicitly
Blank fields can mean different things:
- field not published;
- profile private;
- profile unavailable;
- collection failed;
- field excluded from the project;
- value present but failed validation.
Add fields such as:
profile_status
enrichment_status
contact_field_status
qa_noteExample values:
public_profile_processed
private_profile_identity_only
profile_unavailable
field_not_published
collection_retry_exhausted
excluded_from_scopeThis prevents a blank email from being mistaken for an extraction error.
Include collection timestamps at the correct grain
The follower membership and public profile may be observed at different times.
Use:
membership_collected_at
profile_collected_at
post_collected_at
comment_collected_atOne global project date is helpful but not sufficient for multi-day or repeated work.
Join comments without duplicating profiles
The post–comment–profile structure should be:
posts.post_id
comments.post_id
comments.commenter_profile_id
profiles.master_profile_idIf a username is the only common field, document that join risk. A stable user ID is preferable where available.
For the practical distinction between follower, commenter, and profile datasets, use the Instagram Audience Data Collection guide.
Add source and QA metadata
Recommended QA fields:
source_account_idsource_post_idcollection_run_idparser_versionprofile_statusrequired_field_missingduplicate_resolution_statusqa_statusqa_note
Recommended delivery-level metrics:
- requested identities;
- collected identities;
- unique profiles;
- cross-account duplicates;
- public profiles processed;
- private or unavailable profiles;
- biography coverage;
- website coverage;
- public email coverage;
- failed records by reason.
Security and privacy by design
Collect only what the approved purpose requires. Avoid inferring sensitive traits from names, photos, biographies, follows, or comments. Separate public contact fields from assumptions about permission to contact.
Instagram's Terms prohibit unauthorized information collection, and privacy obligations can apply even when the original profile is public. Instagram Terms of Use (opens in a new tab)
The Instagram Audience Data Collection guide also defines the collection boundaries used across this cluster. Retention, access, minimization, and downstream use should be reviewed for the project's purpose and applicable jurisdiction.
Frequently asked questions
Should one profile have one row?
Use one row per profile in the master table, but retain separate rows for each source membership and comment relationship.
Is username a permanent identifier?
No. Usernames can change. Use a stable user ID when legitimately available and retain observation history.
Should private profiles remain in the follower table?
An identity and private-status observation may remain if it is part of the approved accessible audience list. Private content should not be collected.
How many bio-link columns should a CSV contain?
Use the agreed maximum based on the source sample. For flexible systems, a separate links table is better.
Should public emails be normalized?
Preserve the raw published value, then add a normalized or validation field if that work is included. Do not present syntax validation as proof of consent, ownership, or deliverability.
Next step
Implement the schema through the multi-account Instagram data pipeline, keeping source membership and profile observations at separate grains. For a custom follower, profile, comment, or database schema, review the Instagram Audience Data Collection offer.
Sources
- Meta for Developers: Instagram Platform (opens in a new tab)
- Meta for Developers: Instagram Comment (opens in a new tab)
- RFC 3339: Date and Time on the Internet (opens in a new tab)
- RFC 4180: Common Format and MIME Type for CSV Files (opens in a new tab)
- W3C PROV Data Model (opens in a new tab)
- Instagram Terms of Use (opens in a new tab)

