惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
News | PayPal Newsroom
IT之家
IT之家
Jina AI
Jina AI
博客园 - 司徒正美
GbyAI
GbyAI
WordPress大学
WordPress大学
B
Blog
大猫的无限游戏
大猫的无限游戏
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
Blog — PlanetScale
Blog — PlanetScale
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Recorded Future
Recorded Future
T
Threat Research - Cisco Blogs
AWS News Blog
AWS News Blog
Latest news
Latest news
宝玉的分享
宝玉的分享
小众软件
小众软件
NISL@THU
NISL@THU
C
CERT Recently Published Vulnerability Notes
The GitHub Blog
The GitHub Blog
P
Privacy & Cybersecurity Law Blog
P
Palo Alto Networks Blog
Spread Privacy
Spread Privacy
Last Week in AI
Last Week in AI
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
P
Proofpoint News Feed
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
量子位
博客园_首页
T
The Exploit Database - CXSecurity.com
The Cloudflare Blog
M
MIT News - Artificial intelligence
H
Help Net Security
Security Archives - TechRepublic
Security Archives - TechRepublic
V2EX - 技术
V2EX - 技术
I
InfoQ
D
Darknet – Hacking Tools, Hacker News & Cyber Security
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
O
OpenAI News
MongoDB | Blog
MongoDB | Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Privacy International News Feed
Microsoft Security Blog
Microsoft Security Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Google DeepMind News
Google DeepMind News
H
Hacker News: Front Page
W
WeLiveSecurity
N
News and Events Feed by Topic

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Equity Crowdfunding Leads: scrape 4,800+ Wefunder founders for $5/1K
Devil Scrapes · 2026-05-31 · via DEV Community

Quick answer: There is no unified API for Wefunder, Republic, or StartEngine. An equity crowdfunding leads scraper collects currently-raising and recently-funded campaign data — founder names, company taglines, raise progress, pre-money valuations — from all three platforms and returns them as one normalized dataset. The Apify Actor below does it for $0.005 per row (~$5.05 per 1,000), with the TLS fingerprinting, proxy rotation, and per-source parsing handled for you.

Wefunder alone lists 4,800+ currently-raising companies — founder names, taglines, raise totals, and pre-money valuations in one JSON payload. Republic has a trending carousel. StartEngine has an XML sitemap of 98 offering slugs. None has a download button; none shares a schema.

If you're a VC scout, an SDR targeting founders, or an analyst tracking what's raising in climate versus fintech, you're opening three browser tabs and copy-pasting. Here's what it takes to do that programmatically — and how I compressed it to one API call.

What is equity crowdfunding? 🔎

Equity crowdfunding under Regulation CF lets any US startup raise up to $5 million per year from the general public — not just accredited investors. The three dominant platforms are Wefunder (largest by volume), Republic (curated campaigns), and StartEngine (heavy on CPG and consumer brands).

Each platform requires issuers to file a Form C with the SEC before opening a round, so every active campaign has a verified company name, founding team, financial disclosures, and valuation on public record. That's the dataset: comprehensive, legally disclosed, and — until this Actor — only accessible by visiting three separate sites with three separate UX patterns.

Does Wefunder have an API? 📡

No public API. As of 2026, none of Wefunder, Republic, or StartEngine publishes an official data API or bulk export. Wefunder's SPA calls an internal JSON endpoint (/-/companies/explore) returning full campaign payloads — but it's undocumented, inspects your TLS fingerprint, and sits behind Cloudflare. Republic's backend GraphQL at api.republic.com rejects unauthenticated POSTs from datacenter IPs. StartEngine's offering detail pages require clearing a JavaScript-gated challenge first.

This is exactly why a hosted Actor earns its keep over a three-line requests snippet.

What the data looks like

Each row is a flat, typed record. A real one — RISE Robotics on Wefunder as of 2026-05-16:

{
  "source": "wefunder",
  "campaign_slug": "riserobotics",
  "company_name": "RISE Robotics",
  "tagline": "Electrifying heavy machines",
  "industry": null,
  "location": "MA",
  "founders": ["Hiten Sonpal"],
  "website_url": null,
  "target_amount_usd": null,
  "raised_amount_usd": 17448682.0,
  "num_investors": 417,
  "valuation_usd": 62100000.0,
  "revenue_usd": null,
  "funding_stage": "raising",
  "campaign_url": "https://wefunder.com/riserobotics",
  "scraped_at": "2026-05-16T13:40:00.000Z"
}

Enter fullscreen mode Exit fullscreen mode

Sixteen fields, Pydantic-validated before they hit your dataset. valuation_usd comes from Wefunder's terms.nb shorthand ("$62.1M"), parsed into a float automatically. Republic and StartEngine rows land with the same shape; monetary fields are null there because that data is client-rendered (v2 plan — more below).

The naive approach (and why it falls apart) 🔧

The obvious move: open DevTools, find the XHR, replay it with requests.get(). It breaks fast, for a different reason on each platform.

Wefunder. The /-/companies/explore endpoint checks your TLS ClientHello fingerprint before it answers. Python's stdlib ssl and httpx look nothing like a real browser — the JA3/JA4 fingerprint reads as a script, and you hit a Cloudflare challenge before the JSON loads. We run curl-cffi with impersonate="chrome131", which replays the full Chrome 131 TLS handshake, ALPN extension order, and HTTP/2 SETTINGS frame, so at the TLS layer the connection is a browser.

Republic. The republic.com/companies page is SPA-rendered; the SSR shell carries only a ~10-item carousel of trending campaign links, and the backend GraphQL at api.republic.com rejects unauthenticated POSTs from datacenter IPs. We thread Apify residential proxies on every request so the connection arrives from a residential exit.

StartEngine. Their explore page is fully client-rendered. sitemap-private-offerings.xml carries the active slug list (98 entries as of 2026-05-16) — the only unauthenticated surface; detail pages return a bot-challenge body to non-browser clients. v1 emits slug + company name from the sitemap; Camoufox full-render is planned for v2.

We retry with exponential backoff (base 2 s, doubling, capped at 30 s, max 5 attempts) and honour Retry-After. On 429 or 503 we rotate the proxy session ID — fresh exit IP, fresh cookie jar. Partial success surfaces as an explicit status message; we never return an empty dataset under a green status. One source failing does not kill the run; all three failing exits non-zero with a clear error.

The Actor 🛠️

Equity Crowdfunding Leads on the Apify Store.

Open it in the Apify Console and click Start, or call it with the apify-client Python SDK:

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("DevilScrapes/equity-crowdfunding-leads").call(
    run_input={
        "sources": ["wefunder"],
        "maxPerSource": 200,
        "statusFilter": "active",
        "industryFilter": "fintech",
        "useProxy": True,
    }
)

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["company_name"], item["raised_amount_usd"], item["founders"])

Enter fullscreen mode Exit fullscreen mode

The key input parameters:

  • sources — any combination of wefunder, republic, startengine, or empty (= all three). Default: all three.
  • maxPerSource — hard cap per platform, 1–500. Default: 50.
  • statusFilter"active", "funded", or "all". Wefunder-native; Republic and StartEngine emit currently-listed slugs regardless.
  • industryFilter — optional case-insensitive substring matched against tagline or industry. Pass "climate" for climate-tech campaigns.
  • useProxy — default true. Wefunder and Republic fingerprint datacenter IPs and block plain exits; leave it on.

What you'd actually use this for 💡

Four scenarios from the README and spec:

VC scout pipeline. Schedule a weekly Wefunder-only run, pull all active campaigns, join on founders[], enrich with LinkedIn. A live feed of sub-Series-A founders without waiting for Crunchbase. Scope it with industryFilter: "robotics" for your thesis vertical.

SDR founder outreach. Founders in active crowdfunding campaigns are fundraising — and buying. Filter by statusFilter: "active" and industryFilter: "fintech", drop founders[] into Apollo or Clay, and reach them while they're in motion.

Crowdfunding analytics. Schedule daily runs, persist to BigQuery or S3, and track raised_amount_usd trajectories. Wefunder publishes pre-money valuations Crunchbase never sees — the valuation_usd distribution by sector is a clean dataset for a leaderboard or research report.

Form C deep dives. This Actor surfaces campaign_slug and campaign_url. sec-edgar-filings-scraper (sibling Actor) takes it from there — issuer CIK on EDGAR, Form C / Form C-AR PDFs, audited revenue, SAFE terms. Two Actors, one Reg CF pipeline.

Pricing — exact numbers 💰

Pay-per-event. You pay for rows you receive, nothing for rows that don't come back.

Event Price
Actor start (once per run) $0.05
Per campaign row emitted $0.005
Run size Cost
50 rows (default, all 3 sources) $0.30
150 rows (50/source × 3) $0.80
1,000 rows $5.05
5,000 rows $25.05
10,000 rows $50.05

For context, the nearest alternative — scraping Crunchbase via a third-party Apify Actor — typically runs around $30 per 1,000 rows, while covering fewer than 30% of Wefunder campaigns and zero Republic trending campaigns. This Actor is roughly 6× cheaper and sources from the campaigns directly, not from a derived database. Apify's $5 free trial credit covers your first ~990 rows with no credit card.

The part worth knowing before you build on this 🔍

Wefunder's internal /-/companies/explore endpoint is the same one the SPA calls on every page load — unauthenticated, returning full JSON payloads including pre-money valuation encoded as terms.nb dollar shorthand ("$62.1M", "$700K", "$1.2B"). This Actor parses that shorthand with multipliers K=1e3, M=1e6, B=1e9; malformed values emit null rather than crashing.

The design point worth knowing: the scraper doesn't infer valuations — it reads the exact payload the website reads and converts the display string to a typed float. The Pydantic v2 ResultRow model enforces the schema on every row before write, so type surprises are caught at write time, not at analysis time.

Limitations (the honest list) 🚧

  • Republic and StartEngine return sparse data in v1. Republic surfaces ~10 trending campaign slugs per run from the SSR shell; StartEngine emits slug + company name from the public sitemap. On both, raised amount, valuation, and investor count are client-rendered and stay null. For the richest rows, run Wefunder-only (sources: ["wefunder"]).
  • No historical archive. Every run is a fresh snapshot of currently-listed campaigns. Schedule runs and export to your own storage; Apify's default run-scoped storage is purged after 7 days on the free plan.
  • Status filter is Wefunder-native. funded and all only meaningfully change Wefunder results; Republic and StartEngine always emit their current listing surface regardless.
  • No investor identity data. Who invested and at what amount is private. This Actor emits only public-facing campaign metadata.
  • No SEC EDGAR Form C parsing. Revenue, expenses, share count, and SAFE terms from Form C filings are in scope for sec-edgar-filings-scraper, not this Actor.

FAQ

Is scraping Wefunder, Republic, and StartEngine legal?
All three host public-facing marketing pages built to attract investors. This Actor reads only what the public UI exposes — no authentication is bypassed, no private investor data is collected, and the request rate stays well under a human browsing the site. Form C filings are SEC-required public disclosures. Check your own jurisdiction and use case; nothing here is legal advice.

Does Wefunder, Republic, or StartEngine have an official API I should use instead?
No. As of 2026, none of the three offers a public data API or bulk export endpoint. Wefunder operates an internal JSON endpoint the SPA uses; Republic and StartEngine surface their data via their web UIs (or, for StartEngine, a sitemap).

Can I export the dataset to Google Sheets or a data warehouse?
Yes — export CSV, JSON, Excel, or XML from the Apify Console Export button after the run, webhook the dataset on ACTOR.RUN.SUCCEEDED into Make, Zapier, or n8n, or pull it via the Apify API.

Why does the Actor cost less than Crunchbase scrapers?
Different source, lower extraction cost. Crunchbase scraping hits a richer, more heavily defended site with far more fields. This Actor targets three smaller platforms and returns a narrower, well-defined schema. The 6× difference reflects the actual engineering complexity.

Try it

Live on the Apify Store: apify.com/DevilScrapes/equity-crowdfunding-leads.

Free $5 trial credit, no credit card. Run the defaults and you'll have 150 equity-crowdfunding leads across all three platforms in a couple of minutes. Need a fourth platform (NextSeed, MicroVentures), a field you wish was populated, or a parser that broke after a site restructure? Drop it in the comments. The devil's in the data; I ship based on what people actually find there.


Further reading:


Built by Devil Scrapes — Apify Actors for builders who want the data, not the drama. Pay-per-event, honest pricing, no junk fields. 😈