Blog

Skills analysis

What AI and data employers actually ask for in 411 open jobs

Python and SQL are the common core. Frameworks matter, but the listings point to a broader production stack spanning cloud, data platforms, and deployment.

8 min readJobToaster Research
411
AI & data jobs

Across 71 employers

267
mention Python

The most common named skill

39.2%
remote

161 roles explicitly remote

The common core is Python plus SQL

Python appears in 267 of 411 active Data & AI listings, while SQL appears in 195. That pairing spans data science, machine learning, analytics, and data engineering because it connects modeling with the less glamorous work of retrieving, validating, and transforming production data.

PyTorch is the leading named ML framework at 87 mentions, ahead of TensorFlow at 54. But platform tools are nearly as visible: Spark appears 81 times, AWS 66, Snowflake 61, Google Cloud 47, Databricks 43, and Kubernetes 37. Employers are not only hiring people who can train a model. They need people who can place data and models inside reliable systems.

Most frequently mentioned skills in 411 active Data & AI jobs
SkillJob descriptions mentioning it
Python267
SQL195
PyTorch87
Spark81
AWS66
Snowflake61
TensorFlow54
Google Cloud47
Databricks43
Kubernetes37
Azure29

Tool counts are signals, not a shopping list

A skill mention can be a requirement, preference, team-context reference, or description of the existing stack. The counts are also not mutually exclusive; one posting may mention Python, SQL, AWS, Spark, and Kubernetes. Treat frequency as evidence of interoperability, not a demand to learn eleven tools at once.

A coherent candidate profile is stronger than a checklist. For an applied scientist, that might be Python, statistics, PyTorch, evaluation, and one cloud environment. For a data engineer, it might be SQL, Python, Spark, orchestration, and warehouse design. For an analytics engineer, depth in SQL, modeling, data quality, and stakeholder communication can matter more than a deep-learning framework.

The market is senior, but not exclusively so

Mid-level roles are the largest group at 149, followed by Senior at 136 and Staff at 102. Only one listing was explicitly classified as an internship in this snapshot. The remaining roles include Lead, Principal, Director, and Executive positions.

This distribution rewards proof of production judgment: how data quality was measured, how a model was evaluated after launch, which trade-offs constrained an architecture, and how an analysis changed a decision. A notebook can show technique; a strong portfolio explains the system and consequence around it.

Seniority distribution in active Data & AI jobs
LevelOpen roles
Mid149
Senior136
Staff102
Lead12
Principal9
Executive1
Director1
Internship1

Official projections support the long-term demand signal

The US Bureau of Labor Statistics projects data scientist employment to grow 34% from 2024 to 2034, with about 23,400 openings per year on average. It reports a May 2024 median annual wage of $112,590 and notes that some employers prefer advanced degrees.

The projection is broader than the jobs in this analysis, but the direction is consistent. Data work is spreading beyond research teams into product, security, operations, finance, and infrastructure. That helps explain why production and communication skills recur alongside statistical methods.

A practical learning order from the listings

Start with the common substrate, then specialize. Python and SQL create the widest surface area. Add statistics and experimental reasoning for data science, distributed processing and modeling for data engineering, or one deep-learning framework plus evaluation for ML engineering. Only then add the cloud and deployment tools that fit the target role.

The final differentiator is domain evidence. A candidate who can explain fraud, search, developer tools, healthcare, or marketplace behavior often has a stronger signal than one more generic framework badge. Context turns tools into judgment.

  • Foundation: Python, SQL, data modeling, and version control.
  • Role depth: statistics, distributed systems, or model development.
  • Production layer: one cloud, observability, testing, and deployment.
  • Evidence: a project with a clear user, constraint, metric, and post-launch result.

Sources

  1. 01US Bureau of Labor Statistics: Data Scientists
  2. 02Eurostat: Labour market demand in online job advertisements