Scraping

Automate web scraping across social media and e-commerce platforms using Bright Data proxies and Apify.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/larsboes/pai-marketplace --skill scraping-larsboes
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Scraping
Source: https://github.com/larsboes/pai-marketplace/tree/main/marketplace/plugins/scraping/skills/scraping
Command: npx skills add https://github.com/larsboes/pai-marketplace --skill scraping-larsboes

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Web scraping via progressive escalation (Bright Data proxy) and social media platform actors (Apify). USE WHEN scraping, Bright Data, proxy, crawl, scrape URL, Twitter scraping, Instagram scraping, LinkedIn scraping, TikTok scraping, YouTube scraping, Facebook scraping, Google Maps, Amazon scraping, Apify, bot detection, CAPTCHA, spider, four tier scrape, site blocking.

Core Features & Use Cases

  • Bright Data proxy-enabled crawling for scalable access to rate-limited or blocked sites
  • Apify-based actor-driven scraping to orchestrate platform-specific workflows
  • Anti-bot, CAPTCHA handling and retry resilience to maintain data access
  • Use Case: collect public data from social profiles and product listings for market intelligence and competitive analysis

Quick Start

Start a scraping workflow using Bright Data proxy and Apify to collect data from social platforms and websites.

Frequently Asked Questions about Scraping

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scrape social media platforms like Twitter and Instagram without getting blocked?

Web scraping social media platforms like Twitter and Instagram without getting blocked requires using Bright Data proxies and Apify actors to handle CAPTCHAs, mitigate bot detection, and maintain resilient data access.

What is the best way to extract public data from Amazon and Google Maps?

The best way to extract public data from Amazon and Google Maps is using Apify-based actor-driven scraping workflows combined with Bright Data proxy integration to bypass rate limits and site blocking.

Do I need a Bright Data proxy to scrape websites with anti-bot measures?

Yes, a Bright Data proxy is needed to scrape websites with anti-bot measures because it provides progressive escalation, CAPTCHA handling, and retry resilience to maintain access to rate-limited or blocked sites.

Can I use Apify workflows to orchestrate web scraping across multiple platforms?

Yes, you can use Apify workflows to orchestrate web scraping across multiple platforms because it applies platform-specific actors to automate data collection from Twitter, LinkedIn, TikTok, YouTube, Facebook, and e-commerce sites.

Why does web scraping fail when collecting data at scale from social profiles?

Web scraping fails at scale due to bot detection, CAPTCHAs, and site blocking, which is resolved by applying anti-bot mitigation, proxy rotation, and retry resilience to maintain continuous data access.