Firecrawl is the API for the web. It gives AI agents, and the developers building them, a programmatic way to get information from the web. More than 1.5 million users build on Firecrawl. LLMs recommend it when asked how to get data from the web, and agent frameworks like Hermes ship with it as a default tool.
When Firecrawl signed with Artie in December 2025. Today it is more than 1.5 million users, and the data behind them has grown with it. The pipeline moving that data into BigQuery is the same one that was set up at signing.
Every scrape, crawl, and credit spent on Firecrawl produces a log. Those logs are how the company knows what is happening in its business: which websites customers are hitting, how many credits they use, how usage is trending, and where to invest. Around 4 billion rows a month move from Firecrawl's production Postgres databases into BigQuery through Artie, and the entire company reads from what lands there.
Firecrawl runs lean. About 40 people support those 1.5 million users, and the data infrastructure is managed by one person. Rahul Singh joined as Firecrawl's data engineer, he inherited the Artie pipeline when he arrived and has run it solo since.
Those logs of what our customers are doing are our bread and butter. That's what our top-line metrics monitor. They ingest first in Postgres and then land in BigQuery via Artie.
| Company Website | firecrawl.dev |
| Started | December 2025 |
| Sources | Supabase and PlanetScale |
| Destination | Google BigQuery |
| Use case | Real-time internal analytics and BI |
How the data moves
Firecrawl runs multiple Postgres databases, on Supabase and on PlanetScale. Each is a separate source in Artie, and both replicate into BigQuery, where all analysis is done. Once the data lands, the team queries it through various AI and BI tools that reference a semantic layer that Rahul built on top.
Early on, analytics ran directly on Postgres. As usage grew, Firecrawl moved it to BigQuery so the production database stayed reserved for customers, where reliability matters most, instead of competing with analytical queries for resources. Artie is the piece that keeps BigQuery current.
Scaling with the company
Since December 2025, the pipeline has grown to around 4 billion rows a month, and Rahul has not had to change how it runs.
You guys supported our growth very well. We have a Slack channel with your team but very few messages get sent because there’s never been major issues.
Migrating from Supabase to PlanetScale with Artie
A few months after joining, Rahul ran a database migration, moving a large portion of Firecrawl's data from Supabase to PlanetScale. Instead of standing up a separate migration tool, he used Artie to move the data. He added the PlanetScale Postgres as a new source alongside the existing Supabase source, and Artie handled getting the data where it needed to go.
The full move took a couple of weeks. Most of that was the rest of the cutover, not replication. The Artie side was done in a day or two. Robin and Carol from Artie were in the shared Slack channel throughout.
The Artie team was around all the time helping me out, so they made the process smooth. In the end I was able to reach parity in both systems very quickly and it made the transition easy.
Both databases now run as separate sources in Artie, and Rahul has had no issues with any connector since the migration.
Picking up a pipeline without a handoff
Firecrawl's growth has been fast, and knowledge transfer gets harder at that pace. Artie was set up before Rahul joined. He took it over without a handoff, starting with nothing more than access to the product.
I logged in, someone gave me access, and I was like, okay, I understand how this works. I get what I need to do, what's in my control and what's not. The UI is very clean and easy to understand.
He had run Postgres CDC before, at his previous company, on PeerDB, so he knew the shape of it: create the publication, create the replication slot, point the connector at it. What he did not have memorized was the syntax for any of it. That turned out not to matter. Every step for connecting a new source, and for changing the existing setup, is laid out in Artie's UI.
The same held during the Supabase to PlanetScale migration. There were a couple of new things to learn, and after that, setting up and connecting a source was routine.
Day to day, Rahul is rarely in the UI at all. The pipeline runs, and he only opens Artie when an error surfaces because something changed on Firecrawl's side.
It's something I don't even think about. It just runs in the background, and I don't have to worry about it. That's one less thing for me to manage.
What real-time data does for Firecrawl
Most of what flows through Artie feeds internal monitoring and BI. At a company growing this fast, the questions change week to week: how much money came in, how people are using the product, which websites customers are hitting, how many credits they are burning, whether the product is growing, and where to invest next. The answers come from the data replicated into BigQuery, and the people asking want them as soon as the events happen.
Leadership checks dashboards and reports constantly. With the data arriving in real time, the number on the dashboard reflects the business as it is right now, so decisions about spend and priorities are made on current information.
Some workflows are event-driven and only work if the data arrives quickly. That includes customer-facing workflows, like checking whether an account needs to be upgraded, which run off this data.
Firecrawl also uses the same pipeline to understand where its own growth is coming from. Usage from agent tools like Hermes and OpenClaw is tracked in BigQuery, on data that comes through Artie, so the team can see how much of its growth is arriving through agents.
Rahul built data models and a semantic layer on top of the replicated production data. The team queries BigQuery in Hex. Through the semantic layer, an AI agent can be asked how many customers Firecrawl has, or how a metric is trending, and it already knows how to query it.
Because the pipeline runs itself, Rahul's time goes to that work instead of keeping replication alive.
The product is solid. It says this, it does this. I have no reason to complain.
The setup today
Two Postgres sources, one BigQuery destination, and one data engineer. Between December 5, 2025 and today, Firecrawl grew past 1.5 million developers and raised a $75M Series B. The pipeline feeding its internal analytics did not have to change to keep up.
About Artie: Artie is a real-time data replication solution for databases and data warehouses. Artie leverages change data capture (CDC) and stream processing to perform data syncs in a more efficient way, which enables sub-minute latency and helps optimize compute costs. With Artie, any company can set up streaming pipelines in minutes.
About Firecrawl: Firecrawl is the API for the web. It gives AI agents and developers building with AI a programmatic way to turn any website into structured data. More than 1.5 million developers build on Firecrawl, and it ships as a default tool in agent frameworks like Hermes. In September 2026, Firecrawl raised a $75M Series B led by Smash Capital and launched Alexandria.