Data Analysis

Netflix Animation Production Analysis

Analysis of animated shows and films in which Netflix should produce to grow their Animation viewer retention.

Netflix Analysis Thumbnail

Python · SQL · AWS · Pandas


Project Objective:

What problem are you solving?

I want to find data that would show the current trends in Animated films and TV shows to see how Netflix could expand the retention of its Animation fans and potentially find genres or themes that are topical yet are not being fully capitalized on.

How are you solving this problem?

I explored social media platforms such as Twitter and Reddit using an API to explore trends for films and shows, capturing current interests that may stay relevant for the next few years, and explore reviews for Netflix’s current content and see where fans may find disinterest or a lack in quality, found by cross comparing with results from IMDB.

Job Description

Development Analyst is the position I chose with Netflix' animation department. The responsibilities involve budgeting and handling cash flows for films in development, as well as creating monthly reports to summarize the trends, progress, and health over the film development slate. This project is intended to go in depth with current animations and see how Netflix can potentially target more profitable films to produce in its animation department.

Getting Started

For this project, we will be running everything in Jupyter Notebooks. However, I encourage you to run queries in MySQL workbench for exploration before presenting queries here.

We will be using Reddit’s API to find public reviews which we will compare across IMDB and Netflix’s Animation listings.

NOTE: Make sure to download the datasets listed above. Throughout this project, I stored them in an AWS Database in order to access them and query them within this project.

Accessing Reddit API

First, import all essential Python libraries for this project. We may add more along the way.




Next, we will be establishing credentials for accessing the Reddit API:

#Enter the info for app id, api key
app_id = 'APP_ID'
secret_key = 'SECRET_KEY'

# gain authorization to access api
auth = requests.auth.HTTPBasicAuth(app_id, secret_key)

# set user info to variables
username = 'USERNAME'
password = 'PASSWORD'

# create dictionary that will be passed as an argument in auth request
data = {
    'grant_type':'password',
    'username':username,
    'password':password
}
# add headers that will be passed in auth request
headers = {'User-Agent':'Tutorial12/0.0.1'}

# make authorization request
request = requests.post('<https://www.reddit.com/api/v1/access_token>

#Enter the info for app id, api key
app_id = 'APP_ID'
secret_key = 'SECRET_KEY'

# gain authorization to access api
auth = requests.auth.HTTPBasicAuth(app_id, secret_key)

# set user info to variables
username = 'USERNAME'
password = 'PASSWORD'

# create dictionary that will be passed as an argument in auth request
data = {
    'grant_type':'password',
    'username':username,
    'password':password
}
# add headers that will be passed in auth request
headers = {'User-Agent':'Tutorial12/0.0.1'}

# make authorization request
request = requests.post('<https://www.reddit.com/api/v1/access_token>

#Enter the info for app id, api key
app_id = 'APP_ID'
secret_key = 'SECRET_KEY'

# gain authorization to access api
auth = requests.auth.HTTPBasicAuth(app_id, secret_key)

# set user info to variables
username = 'USERNAME'
password = 'PASSWORD'

# create dictionary that will be passed as an argument in auth request
data = {
    'grant_type':'password',
    'username':username,
    'password':password
}
# add headers that will be passed in auth request
headers = {'User-Agent':'Tutorial12/0.0.1'}

# make authorization request
request = requests.post('<https://www.reddit.com/api/v1/access_token>

Printing the “request” should give you <Response [200]> which indicates that we have made connection.

Next, we will check the authorization token from the response:

You should see a JSON output containing:

  • access_token

  • token_type

  • expires_in

  • scope

We will be placing the access_token within a defined variable to pass through the ‘headers’ when requesting the data.

# assign auth token to a variable
token = request.json()['access_token']

# update 'headers' dict to include auth token
    #token will expire after a certain amount of time
    #after expiration, will need to perform auth sequence again
headers['Authorization'] = 'bearer {}'.format(token)

# if process was successful, this get function will return [200]
requests.get('<https://oauth.reddit.com/api/v1/me>

# assign auth token to a variable
token = request.json()['access_token']

# update 'headers' dict to include auth token
    #token will expire after a certain amount of time
    #after expiration, will need to perform auth sequence again
headers['Authorization'] = 'bearer {}'.format(token)

# if process was successful, this get function will return [200]
requests.get('<https://oauth.reddit.com/api/v1/me>

# assign auth token to a variable
token = request.json()['access_token']

# update 'headers' dict to include auth token
    #token will expire after a certain amount of time
    #after expiration, will need to perform auth sequence again
headers['Authorization'] = 'bearer {}'.format(token)

# if process was successful, this get function will return [200]
requests.get('<https://oauth.reddit.com/api/v1/me>

Now we will request the data from r/Netflix with our new defined ‘headers’.

api = '<https://oauth.reddit.com>

api = '<https://oauth.reddit.com>

api = '<https://oauth.reddit.com>

We will store the JSON results, and extract only the ‘children’ of the data variables, to later store only the essentials in our own dictionary.

netflix_reddit = res.json()

netflix_reddit = netflix_reddit['data']['children']
netflix_reddit = res.json()

netflix_reddit = netflix_reddit['data']['children']
netflix_reddit = res.json()

netflix_reddit = netflix_reddit['data']['children']



We can now begin storing the data from netflix_reddit to our newly created netflix_data dictionary.




Lastly, store the results in a dataframe and save it as an CSV to query in MySQL.




When I conducted this project, I had stored the data in AWS:





Explorative Queries

Query #1:

This query shows all the current data that is present within my netflix_titles table so that I could refer back to the original table when combing through the data.




Query #2:

I saw that there were some records with no countries recorded, so I decided to sort through all the records with nothing in the country column which appears to be 815 records in total.




Query #3:

I saw that there were some records with no date in date_added recorded, so I decided to sort through all the records with nothing in the date_added column which appears to be 25 records in total.




Query #4:

I saw that there were records that oddly dated back way too far in the past to have a correct release year recorded, so I filtered release_year by less than 1800 since most films would go past then, and I ended up with a lot of incorrect years, or no years at all.




Query #5:

Many of the records had data that was shifted around and would not make sense for the column that they belonged to. The most obvious error was when countries were scattered in the table, so I filtered through almost every column to find the ones with incorrect inputs, counting at 4863 rows with errors.




Query #6

The duration column had a lot of inconsistent time signatures which made it hard for me to use it for time ranking, so I took a look at the rows that use seasons instead of minutes. This was shown not only as a differentiator for movie and shows, but within shows it would also be inconsistent.




Query #7

Here, I was verifying to see if the records were being messed up for any other reason or if they actually had errors, so I sorted through to this example record which shows no country and the season duration.




Query #8

I wanted to sort through my IMDB data and see what the errors were. I saw that many errors came from the title columns so i searched through and saw many of them were barely titles at all and would make it more difficult for me to use those tables.




Query #9

I saw that there was an issue with records that were categorized as shorts where they would have no genre and the startYear/ endYear columns would have unusable recorded dates.




Query #10

There wasn't too many discrepancies in the recorded data from the reddit API, but some posts would not have a body which hinders the whole purpose of recording them.





Questions & Analysis

Question #1:

What are some of the most well-received movies recorded on IMDB?

Business Justification: By knowing what movies have higher attention from viewers, we could see which ones Netflix may want to sign contracts with.

SQL features: I used JOIN functions to connect the title and ratings IMDB tables.




Recommendation:

If possible, it would be in Netflix's best interests to increase their animation content by signing contracts for movies and shows such as Cars, Family Guy, Corpse Bride, Futurama, Coraline and more which all receive more than 100,000 positive votes from fans. This would help to retain its subscriptions from animation enthusiasts and families on their platform.

Question #2

Based on the top movies from IMDB, what animated content already exists on Netflix's platform?

Business Justification: If we know which movies and shows have the best received viewing from fans, we could look into sustaining some of our current contracts and potentially renew them.

SQL features: I used a JOIN function to connect the imdb title and rating tables, as well as the netflix title table on the titles.




Recommendation:

Based on the top 10 animated content we hold, we should continue to hold contracts with these titles and make them more noticeable to viewers. In addition, we should attempt to contract with the content that is noted on IMDB that we may not already have.

Question #3

What are people currently saying about our animated content at Netflix?

Business Justification: We want to understand where our clients may have criticism in order to update our platform towards their concerns and needs

SQL features: I unfortunately could not link this with any of the other tables, this was difficult to query for additional info.




Recommendation:

Based on the comments provided by our clients, they express more concern on the loss of our content as well as technical issues that may hinder their user experience. It is in our best interest to continue to withhold our current contracts to maintain a minimal subscription base as well as amplify our efforts to the settings and tools that users have when watching their favorite shows and movies.

Issues:

Initially the goal was to review Netflix Animation listings using Netflix’s API with the IMDB sources, and compare findings with the general consensus of reviews found from Twitter and Reddit, using both their APIs.

This way, we could present to stakeholders, what kinds of animated films and shows should Netflix look towards producing.

However, this project came with several caveats.

  1. Netflix API

I started this project expecting to use the Netflix API to gather data on revenue and gross profit from certain films, but halfway through this project, I found out that Netflix had shut down their API a few years ago. So I needed to resort to a Kaggle public dataset as a demonstrative source for this kind of information.

  1. Twitter API

Initially, I would have liked to implement my Twitter API, however, I would have needed to request for elevated permissions in order to properly use it in a timely manner. It was not accessible with how new my account was.

  1. Reddit Reviews

Much of the data did not span large enough to properly compare to what viewers thought of the current state of Netflix’s animation listings. Much of was arbitrary topics and was difficult to flush out.

  1. Corrupted Data

In several instances, I kept finding data that would shift and redistribute into incorrect fields. This should have been an obvious sign to clean and correct the data, but with minimal time constraint, I was unable to prioritize this essential task, and had to resort to data that was in-tact.


Final Thoughts

This was a fun project to try and tackle, however, it did come with several challenges that were not originally predictable. If I were to go through this project again, I would have gone through this project differently:

  1. Data Sources

Due to short time constraints, I was unable to thoroughly research and take my time with this as I should have. Data sources should have been chosen much more deliberately from the start. I was ambitious to try to use social media APIs to analyze public commentary on movies and shows, but I was unaware of the difficulties I would face with user permissions and technical setup.

And without the Netflix API, relying on Kaggle for public datasets was more of a crutch than a reliable leg to this project. IMDB sources was a good hand to have throughout this project, but didn’t always have consistent data. I had to rely on a free-tier version which had limitations. But for learning purposes, sufficed.

  1. Changes

If I were to redo this project, I would have removed any use of social media APIs and looked for any new APIs that may be accessible for the many new streaming services that we have today. This project was conducted around 2022, so potentially, more sources may be accessible nowadays.

Furthermore, I would have spent a significant more time on data cleaning to ensure more data is usable throughout this project, instead of focusing on configuring social media APIs.

I would have also liked to have been able to produce visuals to help guide stakeholders of key films and shows to reference for optimal content to grow Netflix’s animation viewer retention.

However, on the subject of social media; it would have been interesting to explore YouTube and TikTok data on viewer interest in Shorts, now that Disney+ and other streaming services are working to create algorithmic features for short-form content.

As of 2026, Netflix has already grown their animation department and has been working to continuously expand their animation listings with a larger following of fans on their platform.



Thank you for reading my findings, I would be happy to share any deeper technical insights or answer any questions you may have.

Please feel free to reach out to me either on LinkedIn or by submitting a contact form here.

Data Analyst transforming data into clear insights and practical solutions.

Sacramento, CA · in-office - hybrid/remote

Sacramento, CA · in-office

- hybrid/remote

© 2026 Andy. All rights reserved.

Turning Data into Insights and Action