Get new posts by email

New posts on data pipelines, scraping, and market-data ML — straight to your inbox.

It's completely free, and you can unsubscribe anytime.

By subscribing you agree to Substack's Terms of Use, its Privacy Policy and its Information collection notice.

Substack

Vishnujan NarayananData pipelines,from crawlers to APIs

I build the crawlers and scheduled jobs that pull data in, the tests that ensure the data is valid, and the APIs and apps that serve it.

scroll to see how I work
Writing

Blogs

Notes from the build — deep dives on data pipelines, applied ML, and the engineering behind shipping systems that hold up in production.

Testing

How to Test a Data Pipeline: What Tests Miss

A green suite is evidence about your code, not your data. The cheap tests that catch what it misses.

Read →
Data engineering

Data Ingestion Bugs and How to Catch Them

Five ways ingestion breaks, with the guard for each. Eight of about 30 bugs threw no error at all.

Read →
Web scraping

How to Scrape a Site That Paginates by Date

Find the undocumented window cap, resume from behind, and dedup on a key that cannot fail silently.

Read →
LLMs

Local LLM vs API: When to Run Your Own Model

What fits in 6GB, what it costs in latency, and the schema trick that makes a small model usable.

Read →
ETL

How to Make an ETL Pipeline Safe to Rerun

A rerun should repair data, not corrupt it. Four ways that rule breaks, and a six-point checklist.

Read →
APIs

HTTP 429: What It Means and How to Retry

One status code, three meanings — wait, stop, or out of quota. How to tell them apart in code.

Read →
Quant / ML

Can You Predict Stock Prices Minute by Minute?

Four checks that decide whether a short-horizon model has an edge. One real attempt failed all four.

Read →
Web scraping

Why Your Scraper Gets Blocked and How to Fix It

Four layers block you: headers, TLS fingerprint, IP, behaviour. How to tell which one, and the fix for each.

Read →
Web scraping

Playwright vs Selenium: Which to Use and When

Locators, saved sessions and pinned browsers — plus the 52 fixed sleeps that show what auto-waiting does not fix.

Read →
15 rows in set (0.001 sec)

The Product

QuickForge

Fresh postings.Instant alerts.Tailored résumé, every time.

QuickForge monitors job listings in real time, scores every posting against your profile, and delivers a tailored résumé for each match — send quality applications before quantity shows up.

QuickForge product preview

No callbacks?
Your resume isn’t talking their language.

Recruiters get 200+ per role. Generic gets filtered before a human looks. Your experience isn’t the problem.

The fix

The bot reads the JD and selects the exact bullets, title and skill framing that match it. Every application speaks that role’s specific language. Automatically. Every time.

Applicant #847.
Window closed 3 hours ago.

Recruiters stop reading after the first 50. The window is under an hour from posting. You can’t watch three boards while that window opens and shuts.

The fix

The bot monitors Indeed, Glassdoor and LinkedIn continuously. The moment a match posts, you get a Telegram ping with the apply link and a tailored résumé ready to go. First wave, not after it.

Tweaking the same
résumé. Again.

10 roles a day is 10 tailoring sessions. You’re spending more time in Word than applying. And the rushed ones always look it.

The fix

You write your bullets once. The bot selects the right ones per role, every time. Your master profile becomes every tailored résumé you’ll ever need.

In development · Lab 01

Coming soon

PythonFastAPIPostgres + pgvectorPlaywrightTelegram