Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

Instagram Public Profile Data Schema: Fields, Links, and Provenance

Engineering2026-07-1811 minHira Arif

Design a reliable schema for Instagram usernames, biographies, public metrics, bio links, source membership, comments, and collection-quality fields.

Short answer

A reliable Instagram public-profile schema separates identity, source membership, raw profile observations, parsed biography elements, public contact fields, content relationships, and QA metadata. Do not place every field into one undocumented “lead” row: usernames change, biographies contain free text, one profile can appear under several source accounts, and contact values can be missing or ambiguous.

Four questions every value should answer

The schema should answer four questions for every value:

  1. Which Instagram identity does it describe?
  2. Which source account or post placed the identity in scope?
  3. Was the value published directly, parsed, or normalized?
  4. When was it collected?

The four data grains

Instagram audience projects commonly contain four different grains.

Source account

One row per approved account whose followers, following, posts, or comments are in scope.

Profile

One row per resolved Instagram identity.

Membership or appearance

One row per relationship between a profile and a source account, post, or interaction type.

Content or interaction

One row per post, comment, reply, or other approved public interaction.

Mixing these grains causes duplication. If a profile wrote ten comments, a flat file may repeat its biography, website, and follower count ten times. If that profile also follows three source accounts, the same fields may be repeated again.

Recommended relational model

source_accounts

text
source_account_id
source_username
source_profile_url
source_category
population_type
displayed_count
requested_limit
collection_started_at
collection_completed_at
collection_status
status_note

profiles

text
master_profile_id
instagram_user_id
username_raw
username_current
display_name_raw
profile_url
is_private
is_verified
profile_status
profile_collected_at

profile_observations

text
master_profile_id
observed_at
biography_raw
followers_count
following_count
posts_count
business_category_raw
external_url_raw
public_email_raw
public_phone_raw

audience_memberships

text
source_account_id
master_profile_id
relationship_type
observed_at

Possible relationship_type values include:

text
follower
following
commenter
replier
liker
mentioned_profile
tagged_profile
client_provided

profile_links

text
master_profile_id
link_position
link_url_raw
link_url_normalized
link_domain
link_type
observed_at

posts

text
post_id
source_account_id
post_url
post_type
caption_raw
published_at
public_like_count
public_comment_count
collected_at

comments

text
comment_id
post_id
commenter_profile_id
parent_comment_id
comment_text_raw
commented_at
public_like_count
collected_at

This model supports follower exports, overlap analysis, comment research, and repeated profile snapshots without overwriting history.

A practical flat-file schema

Not every project needs a database. A well-defined Excel or CSV delivery can still be reliable.

Recommended master fields:

FieldMeaning
master_profile_idClient-side stable identifier
instagram_user_idPlatform identity when available
usernameUsername observed during enrichment
display_name_rawDisplay name exactly as published
biography_rawFull biography before parsing
bio_hashtagsHashtags parsed from biography
bio_mentionsMentions parsed from biography
followers_countPublic count observed during collection
following_countPublic count observed during collection
posts_countPublic count observed during collection
is_privateProfile privacy status observed
is_verifiedVerification status observed
business_category_rawCategory as published
public_email_rawEmail visibly published when available
public_phone_rawPhone visibly published when available
bio_linksPublic links retained in order
source_accountsAccounts that placed the profile in scope
source_countNumber of unique source memberships
profile_urlPublic profile URL
profile_collected_atTime of the profile observation
qa_statusDelivery review status

Preserve raw and parsed biography fields

An Instagram biography is free text. It may contain:

  • names and roles;
  • emoji;
  • hashtags;
  • mentions;
  • location phrases;
  • business claims;
  • links typed as text;
  • line breaks;
  • several languages.

Always retain biography_raw before creating parsed fields.

Example:

text
biography_raw: "Independent studio | Chicago + NYC | @projectname | #slowfashion"
bio_mentions: "projectname"
bio_hashtags: "slowfashion"
location_text_raw: "Chicago + NYC"

Do not silently turn Chicago + NYC into a verified city and state. It may describe markets served, previous residence, travel, or something else. If location is important, define a conservative parsing rule and retain the raw text.

Model multiple bio links properly

A profile may publish:

  • one external URL;
  • a link-in-bio page containing several destinations;
  • a booking link;
  • a portfolio;
  • an online store;
  • a newsletter;
  • another social profile.

For a spreadsheet, fixed columns can be practical:

text
bio_link_1
bio_link_2
bio_link_3
bio_link_4
bio_link_5
bio_link_6

This matches Orzaen's six-account entertainment project, which used a 17-field public profile schema with up to six links per profile. See the multi-account entertainment audience case study.

For a database, use a separate profile_links table. Preserve link order and raw URL before normalization.

Do not overwrite raw values with normalized values

Examples:

text
username_raw -> username_normalized
display_name_raw -> display_name_search
public_phone_raw -> public_phone_e164
link_url_raw -> link_url_normalized
business_category_raw -> business_category_mapped

Normalization supports search and comparison, but the raw value is the source observation. If a mapping rule changes later, the record can be reprocessed without revisiting the profile.

Model missing values explicitly

Blank fields can mean different things:

  • field not published;
  • profile private;
  • profile unavailable;
  • collection failed;
  • field excluded from the project;
  • value present but failed validation.

Add fields such as:

text
profile_status
enrichment_status
contact_field_status
qa_note

Example values:

text
public_profile_processed
private_profile_identity_only
profile_unavailable
field_not_published
collection_retry_exhausted
excluded_from_scope

This prevents a blank email from being mistaken for an extraction error.

Include collection timestamps at the correct grain

The follower membership and public profile may be observed at different times.

Use:

text
membership_collected_at
profile_collected_at
post_collected_at
comment_collected_at

One global project date is helpful but not sufficient for multi-day or repeated work.

Join comments without duplicating profiles

The post–comment–profile structure should be:

text
posts.post_id
comments.post_id
comments.commenter_profile_id
profiles.master_profile_id

If a username is the only common field, document that join risk. A stable user ID is preferable where available.

For the practical distinction between follower, commenter, and profile datasets, use the Instagram Audience Data Collection guide.

Add source and QA metadata

Recommended QA fields:

  • source_account_id
  • source_post_id
  • collection_run_id
  • parser_version
  • profile_status
  • required_field_missing
  • duplicate_resolution_status
  • qa_status
  • qa_note

Recommended delivery-level metrics:

  • requested identities;
  • collected identities;
  • unique profiles;
  • cross-account duplicates;
  • public profiles processed;
  • private or unavailable profiles;
  • biography coverage;
  • website coverage;
  • public email coverage;
  • failed records by reason.

Security and privacy by design

Collect only what the approved purpose requires. Avoid inferring sensitive traits from names, photos, biographies, follows, or comments. Separate public contact fields from assumptions about permission to contact.

Instagram's Terms prohibit unauthorized information collection, and privacy obligations can apply even when the original profile is public. Instagram Terms of Use (opens in a new tab)

The Instagram Audience Data Collection guide also defines the collection boundaries used across this cluster. Retention, access, minimization, and downstream use should be reviewed for the project's purpose and applicable jurisdiction.

Frequently asked questions

Should one profile have one row?

Use one row per profile in the master table, but retain separate rows for each source membership and comment relationship.

Is username a permanent identifier?

No. Usernames can change. Use a stable user ID when legitimately available and retain observation history.

Should private profiles remain in the follower table?

An identity and private-status observation may remain if it is part of the approved accessible audience list. Private content should not be collected.

How many bio-link columns should a CSV contain?

Use the agreed maximum based on the source sample. For flexible systems, a separate links table is better.

Should public emails be normalized?

Preserve the raw published value, then add a normalized or validation field if that work is included. Do not present syntax validation as proof of consent, ownership, or deliverability.

Next step

Implement the schema through the multi-account Instagram data pipeline, keeping source membership and profile observations at separate grains. For a custom follower, profile, comment, or database schema, review the Instagram Audience Data Collection offer.

Sources

Tags

Instagram ProfilesData SchemaData EngineeringPublic Data

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review