Convert Figma logo to code with AI

spatie logocrawler

https://spatie.be/docs/crawler

2,814
368
2,814
0

Top Related Projects

🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent

Gospider - Fast web spider written in Go

25,284

Elegant Scraper and Crawler Framework for Golang

63,113

Scrapy, a fast high-level web crawling & scraping framework for Python.

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

5,669

Self-hosted, easily-deployable monitoring and alerts service - like a lightweight PagerDuty

Quick Overview

The spatie/crawler is a PHP package that provides a simple and efficient way to crawl websites. It can be used to extract data from web pages, follow links, and perform various other tasks related to web scraping.

Pros

  • Flexibility: The package offers a flexible and customizable API, allowing developers to easily configure the crawling process to suit their specific needs.
  • Scalability: The crawler can handle large-scale web crawling tasks, making it suitable for a wide range of applications, from content aggregation to SEO analysis.
  • Asynchronous Processing: The package utilizes asynchronous processing, which can significantly improve the performance and efficiency of the crawling process.
  • Robust Error Handling: The crawler provides robust error handling, ensuring that the crawling process can continue even if some pages fail to load or encounter other issues.

Cons

  • Limited Functionality: While the package is powerful, it may not provide all the features and functionality that some users might require for more complex web scraping tasks.
  • Dependency on External Libraries: The package relies on several external libraries, which can increase the complexity of the setup and potentially introduce additional dependencies.
  • Learning Curve: Developers who are new to web scraping or the spatie/crawler package may need to invest some time in understanding the API and how to effectively use the package.
  • Potential Legal Concerns: Web scraping can raise legal concerns, and developers should be aware of the relevant laws and regulations in their jurisdiction before using the package.

Code Examples

Here are a few code examples demonstrating the usage of the spatie/crawler package:

  1. Basic Crawling:
use Spatie\Crawler\Crawler;

Crawler::create()
    ->startUrl('https://example.com')
    ->setCrawlObserver(new MyCrawlObserver())
    ->setDelayBetweenRequests(500)
    ->setMaximumCrawlCount(100)
    ->setMaximumDepth(3)
    ->run();

This code sets up a basic crawling process, starting from the URL https://example.com, using the MyCrawlObserver class to handle the crawling events, and setting some configuration options such as the delay between requests and the maximum crawl count and depth.

  1. Filtering URLs:
use Spatie\Crawler\Crawler;
use Spatie\Crawler\CrawlProfile;

class MyCustomCrawlProfile implements CrawlProfile
{
    public function shouldCrawl(string $url): bool
    {
        return strpos($url, 'example.com') !== false;
    }
}

Crawler::create()
    ->setCrawlProfile(new MyCustomCrawlProfile())
    ->startUrl('https://example.com')
    ->setCrawlObserver(new MyCrawlObserver())
    ->run();

This example demonstrates how to use a custom CrawlProfile to filter the URLs that the crawler should follow. In this case, the MyCustomCrawlProfile class only allows the crawler to follow URLs that contain the string 'example.com'.

  1. Handling Redirects:
use Spatie\Crawler\Crawler;
use Spatie\Crawler\CrawlProfile;

class MyCustomCrawlProfile implements CrawlProfile
{
    public function shouldCrawl(string $url): bool
    {
        return strpos($url, 'example.com') !== false;
    }

    public function shouldCrawlResponse(ResponseInterface $response): bool
    {
        return $response->getStatusCode() !== 301 && $response->getStatusCode() !== 302;
    }
}

Crawler::create()
    ->setCrawlProfile(new MyCustomCrawlProfile())
    ->startUrl('https://example.com')
    ->setCrawlObserver(new MyCrawlObserver())
    ->run();

This example demonstrates how to handle redirects using a custom CrawlProfile. The shouldCrawlResponse method is used to filter out responses with a 301 or 302 status code, which typically indicate a redirect.

Getting Started

Competitor Comparisons

🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent

Pros of Crawler-Detect

  • Crawler-Detect is a lightweight library focused solely on detecting web crawlers, bots, and spiders.
  • It has a comprehensive database of known crawlers, making it effective at identifying a wide range of bots.
  • The library is easy to integrate into web applications and has a simple API.

Cons of Crawler-Detect

  • Crawler-Detect is not as feature-rich as spatie/crawler, which provides more advanced crawling capabilities.
  • The library is primarily focused on detection and does not provide functionality for crawling websites.
  • The database of known crawlers may not be as frequently updated as some other solutions.

Code Comparison

Crawler-Detect:

$crawler = new Crawler();
if ($crawler->isCrawler()) {
    echo "This is a bot!";
}

spatie/crawler:

$crawler = new Crawler();
$crawler->setCrawlObserver(new MyCrawlObserver());
$crawler->startCrawling('https://example.com');

Gospider - Fast web spider written in Go

Pros of gospider

  • Written in Go, potentially offering better performance and concurrency
  • Includes built-in features like subdomain enumeration and JavaScript parsing
  • Designed specifically for web security testing and bug bounty hunting

Cons of gospider

  • Less flexible and customizable compared to crawler
  • May be more complex to use for simple crawling tasks
  • Limited documentation and community support

Code Comparison

crawler (PHP):

Crawler::create()
    ->setCrawlObserver(new MyCrawlObserver)
    ->startCrawling('https://example.com');

gospider (Go):

options := &core.Options{
    Concurrent: 5,
    Depth:      2,
    ParseJS:    true,
}
crawler := core.NewCrawler(options)
crawler.Start("https://example.com")

Both repositories provide web crawling functionality, but they cater to different use cases and ecosystems. crawler is a more general-purpose PHP library with a focus on flexibility and ease of use. It's well-suited for PHP developers who need to integrate crawling capabilities into their applications.

gospider, on the other hand, is a specialized tool written in Go, targeting security researchers and bug bounty hunters. It offers built-in features specific to web security testing, making it more suitable for those scenarios. However, it may be less adaptable for general-purpose crawling tasks compared to crawler.

The choice between the two depends on the specific requirements of your project, your preferred programming language, and whether you need the specialized security-focused features of gospider or the more flexible approach of crawler.

25,284

Elegant Scraper and Crawler Framework for Golang

Pros of colly

  • Written in Go, offering better performance and concurrency
  • Lightweight and easy to use with a simple API
  • Supports distributed scraping out of the box

Cons of colly

  • Less feature-rich compared to crawler's extensive options
  • Limited built-in support for JavaScript rendering
  • Smaller community and ecosystem than PHP-based crawler

Code Comparison

crawler (PHP):

Crawler::create()
    ->setCrawlObserver(new MyCrawlObserver)
    ->startCrawling('https://example.com');

colly (Go):

c := colly.NewCollector()
c.OnHTML("a[href]", func(e *colly.HTMLElement) {
    e.Request.Visit(e.Attr("href"))
})
c.Visit("https://example.com")

Both libraries provide a straightforward way to start crawling a website. crawler offers a more declarative approach with its fluent interface, while colly uses callback functions for handling different aspects of the crawling process. colly's code is more concise but may require more setup for complex scenarios.

63,113

Scrapy, a fast high-level web crawling & scraping framework for Python.

Pros of Scrapy

  • More comprehensive and feature-rich, offering advanced capabilities like item pipelines and middleware
  • Supports multiple output formats (JSON, CSV, XML) out of the box
  • Has a larger community and ecosystem, with more extensions and plugins available

Cons of Scrapy

  • Steeper learning curve due to its complexity and Python-specific concepts
  • Heavier and potentially slower for simple scraping tasks
  • Requires more setup and configuration for basic use cases

Code Comparison

Crawler (PHP):

$crawler = Crawler::create()
    ->setCrawlObserver(new MyCrawlObserver)
    ->startCrawling('https://example.com');

Scrapy (Python):

class MySpider(scrapy.Spider):
    name = 'myspider'
    start_urls = ['https://example.com']

    def parse(self, response):
        # Parsing logic here

Crawler is more straightforward for simple tasks, while Scrapy requires more boilerplate but offers greater flexibility. Crawler is ideal for PHP developers seeking a lightweight solution, whereas Scrapy is better suited for complex, large-scale scraping projects in Python.

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

Pros of Heritrix3

  • More robust and scalable, designed for large-scale web archiving
  • Offers advanced features like deduplication and adaptive crawling
  • Provides a web-based user interface for easier management

Cons of Heritrix3

  • Steeper learning curve and more complex setup
  • Requires more system resources due to its comprehensive nature
  • Less suitable for small-scale or quick crawling tasks

Code Comparison

Crawler (PHP):

Crawler::create()
    ->setCrawlObserver(new MyCrawlObserver)
    ->startCrawling('https://example.com');

Heritrix3 (Java):

CrawlJob job = new CrawlJob();
job.setName("MyJob");
job.addSeed("https://example.com");
job.launch();

Key Differences

  • Crawler is lightweight and easy to integrate into PHP projects
  • Heritrix3 is more feature-rich but requires Java and additional setup
  • Crawler is better for small to medium-scale crawling tasks
  • Heritrix3 excels in large-scale, archival-quality web crawling

Use Cases

  • Choose Crawler for simple web scraping or site auditing in PHP environments
  • Opt for Heritrix3 for comprehensive web archiving or large-scale data collection projects
5,669

Self-hosted, easily-deployable monitoring and alerts service - like a lightweight PagerDuty

Pros of Cabot

  • Comprehensive monitoring solution with web interface and alerting capabilities
  • Supports multiple service checks (HTTP, Jenkins, Graphite metrics)
  • Integrates with various notification channels (email, SMS, Slack)

Cons of Cabot

  • More complex setup and configuration compared to Crawler
  • Focused on monitoring rather than web crawling
  • Less flexibility for custom crawling logic

Code Comparison

Crawler (PHP):

Crawler::create()
    ->setCrawlObserver(new MyCrawlObserver)
    ->startCrawling('https://example.com');

Cabot (Python):

from cabot.cabotapp.models import Service

service = Service.objects.create(
    name="My Website",
    url="https://example.com",
    status_checks=[HttpStatusCheck.objects.create(endpoint="/")]
)

While Crawler is designed for web crawling tasks, Cabot is built for monitoring and alerting. Crawler offers a simpler API for crawling websites, whereas Cabot provides a more comprehensive solution for monitoring various services and metrics. The code examples illustrate the difference in focus: Crawler initializes a crawl job, while Cabot sets up a service to be monitored.

Convert Figma logo designs to code with AI

Visual Copilot

Introducing Visual Copilot: A new AI model to turn Figma designs to high quality code using your components.

Try Visual Copilot

README

Logo for crawler

Crawl the web using PHP

Latest Version on Packagist MIT Licensed Tests Total Downloads

This package provides a powerful, easy to use class to crawl links on a website. Under the hood, Guzzle promises are used to crawl multiple URLs concurrently.

Because the crawler can execute JavaScript, it can crawl JavaScript rendered sites. Under the hood, Chrome and Puppeteer are used to power this feature.

Here's a quick example:

use Spatie\Crawler\Crawler;
use Spatie\Crawler\CrawlResponse;

Crawler::create('https://example.com')
    ->onCrawled(function (string $url, CrawlResponse $response) {
        echo "{$url}: {$response->status()}\n";
    })
    ->start();

Or collect all URLs on a site:

$urls = Crawler::create('https://example.com')
    ->internalOnly()
    ->depth(3)
    ->foundUrls();

You can also test your crawl logic without making real HTTP requests:

Crawler::create('https://example.com')
    ->fake([
        'https://example.com' => '<html><a href="/about">About</a></html>',
        'https://example.com/about' => '<html>About page</html>',
    ])
    ->foundUrls();

If you need to stop a crawl based on external state, you can register a callback that receives the current crawler instance and is checked before scheduling each next request:

use Spatie\Crawler\Crawler;

$shouldStop = false;

Crawler::create('https://example.com')
    ->shouldStopCallback(function (Crawler $crawler) use (&$shouldStop) {
        return $shouldStop;
    })
    ->onCrawled(function (string $url) use (&$shouldStop) {
        $shouldStop = true;
    })
    ->start();

Support us

We invest a lot of resources into creating best in class open source packages. You can support us by buying one of our paid products.

We highly appreciate you sending us a postcard from your hometown, mentioning which of our package(s) you are using. You'll find our address on our contact page. We publish all received postcards on our virtual postcard wall.

Documentation

All documentation is available on our documentation site.

Testing

composer test

Changelog

Please see CHANGELOG for more information on what has changed recently.

Contributing

Please see CONTRIBUTING for details.

Security Vulnerabilities

Please review our security policy on how to report security vulnerabilities.

Credits

License

The MIT License (MIT). Please see License File for more information.