ixnode/php-web-crawler
Composer 安装命令:
composer require ixnode/php-web-crawler
包简介
PHP Web Crawler - This PHP class allows you to crawl recursively a given html page (or a given html file) and collect some data from it.
README 文档
README
This PHP class allows you to crawl recursively a given html page (or a given html file) and collect some data from it. Simply define the url (or a html file) and a set of xpath expressions which should map with the output data object. The final representation will be a php array which can be easily converted into the json format for further processing.
1. Installation
composer require ixnode/php-web-crawler
vendor/bin/php-web-crawler -V
php-web-crawler 0.1.0 (02-24-2024 14:46:26) - Björn Hempel <bjoern@hempel.li>
2. Usage
2.1 PHP Code
use Ixnode\PhpWebCrawler\Output\Field; use Ixnode\PhpWebCrawler\Source\Raw; use Ixnode\PhpWebCrawler\Value\Text; use Ixnode\PhpWebCrawler\Value\XpathTextNode; $rawHtml = <<<HTML <html> <head> <title>Test Page</title> </head> <body> <h1>Test Title</h1> <p>Test Paragraph</p> </body> </html> HTML; $html = new Raw( $rawHtml, new Field('version', new Text('1.0.0')), new Field('title', new XpathTextNode('//h1')), new Field('paragraph', new XpathTextNode('//p')) ); $html->parse()->getJsonStringFormatted(); // See below
2.2 JSON result
{
"version": "1.0.0",
"title": "Test Title",
"paragraph": "Test Paragraph"
}
3. Advanced usage
3.1 Group
PHP Code
use Ixnode\PhpWebCrawler\Output\Field; use Ixnode\PhpWebCrawler\Output\Group; use Ixnode\PhpWebCrawler\Source\Raw; use Ixnode\PhpWebCrawler\Value\XpathTextNode; $rawHtml = <<<HTML <html> <head> <title>Test Page</title> </head> <body> <h1>Test Title</h1> <p class="paragraph-1">Test Paragraph 1</p> <p class="paragraph-2">Test Paragraph 2</p> </body> </html> HTML; $html = new Raw( $rawHtml, new Field('title', new XpathTextNode('/html/head/title')), new Group( 'content', new Group( 'header', new Field('h1', new XpathTextNode('/html/body//h1')), ), new Group( 'text', new Field('p1', new XpathTextNode('/html/body//p[@class="paragraph-1"]')), new Field('p2', new XpathTextNode('/html/body//p[@class="paragraph-2"]')), ) ) ); $html->parse()->getJsonStringFormatted(); // See below
JSON result
{
"title": "Test Page",
"content": {
"header": {
"h1": "Test Title"
},
"text": {
"p1": "Test Paragraph 1",
"p2": "Test Paragraph 2"
}
}
}
3.2 XpathSection
PHP Code
use Ixnode\PhpWebCrawler\Output\Field; use Ixnode\PhpWebCrawler\Output\Group; use Ixnode\PhpWebCrawler\Source\Raw; use Ixnode\PhpWebCrawler\Source\XpathSection; use Ixnode\PhpWebCrawler\Value\XpathTextNode; $rawHtml = <<<HTML <html> <head> <title>Test Page</title> </head> <body> <div class="content"> <h1>Test Title</h1> <p class="paragraph-1">Test Paragraph 1</p> <p class="paragraph-2">Test Paragraph 2</p> </div> </body> </html> HTML; $html = new Raw( $rawHtml, new Field('title', new XpathTextNode('/html/head/title')), new Group( 'content', new XpathSection( '/html/body//div[@class="content"]', new Group( 'header', new Field('h1', new XpathTextNode('./h1')), ), new Group( 'text', new Field('p1', new XpathTextNode('./p[@class="paragraph-1"]')), new Field('p2', new XpathTextNode('./p[@class="paragraph-2"]')), ) ) ) ); $html->parse()->getJsonStringFormatted(); // See below
JSON result
{
"title": "Test Page",
"content": {
"header": {
"h1": "Test Title"
},
"text": {
"p1": "Test Paragraph 1",
"p2": "Test Paragraph 2"
}
}
}
3.3 XpathSection (flat)
PHP Code
use Ixnode\PhpWebCrawler\Output\Field; use Ixnode\PhpWebCrawler\Output\Group; use Ixnode\PhpWebCrawler\Source\Raw; use Ixnode\PhpWebCrawler\Source\XpathSections; use Ixnode\PhpWebCrawler\Value\XpathTextNode; $rawHtml = <<<HTML <html> <head> <title>Test Page</title> </head> <body> <div class="content"> <h1>Test Title</h1> <p class="paragraph-1">Test Paragraph 1</p> <p class="paragraph-2">Test Paragraph 2</p> <ul> <li>Test Item 1</li> <li>Test Item 2</li> </ul> </div> </body> </html> HTML; $html = new Raw( $rawHtml, new Field('title', new XpathTextNode('/html/head/title')), new Group( 'hits', new XpathSections( '/html/body//div[@class="content"]/ul', new XpathTextNode('./li/text()'), ) ) ); $html->parse()->getJsonStringFormatted(); // See below
JSON result
{
"title": "Test Page",
"hits": [
[
"Test Item 1",
"Test Item 2"
]
]
}
3.3 XpathSection (structured)
PHP Code
use Ixnode\PhpWebCrawler\Output\Field; use Ixnode\PhpWebCrawler\Output\Group; use Ixnode\PhpWebCrawler\Source\Raw; use Ixnode\PhpWebCrawler\Source\XpathSections; use Ixnode\PhpWebCrawler\Value\XpathTextNode; $rawHtml = <<<HTML <html> <head> <title>Test Page</title> </head> <body> <div class="content"> <h1>Test Title</h1> <p class="paragraph-1">Test Paragraph 1</p> <p class="paragraph-2">Test Paragraph 2</p> <table> <tbody> <tr> <th>Caption 1</th> <td>Cell 1</td> </tr> <tr> <th>Caption 2</th> <td>Cell 2</td> </tr> </tbody> </table> </div> </body> </html> HTML; $html = new Raw( $rawHtml, new Field('title', new XpathTextNode('/html/head/title')), new Group( 'hits', new XpathSections( '/html/body//div[@class="content"]/table/tbody/tr', new Field('caption', new XpathTextNode('./th/text()')), new Field('content', new XpathTextNode('./td/text()')), ) ) ); $html->parse()->getJsonStringFormatted(); // See below
JSON result
{
"title": "Test Page",
"hits": [
{
"caption": "Caption 1",
"content": "Cell 1"
},
{
"caption": "Caption 2",
"content": "Cell 2"
}
]
}
4. More examples
- examples/converter.php
- examples/group.php
- examples/section.php
- examples/sections-recursive-url.php
- examples/sections.php
- examples/simple-wiki-page.php
5. Development
git clone git@github.com:ixnode/php-web-crawler.git && cd php-web-crawler
composer install
composer test
6. License
This library is licensed under the MIT License - see the LICENSE.md file for details.
ixnode/php-web-crawler 适用场景与选型建议
ixnode/php-web-crawler 是一款 基于 PHP 开发的 Composer 扩展包,目前已累计 46 次下载、GitHub Stars 达 2, 最近一次更新时间为 2024 年 02 月 24 日, 在 PHP 生态内属于活跃度较高的组件。
它主要适用于以下技术方向: 「php」 「json」 「web」 「html」 「array」 「scraper」 等业务场景。在实际项目中,围绕这些方向常见需要落地的问题包括:接口对接、性能调优、并发安全、与既有框架(Laravel / ThinkPHP / Yii / Webman 等)的兼容适配,以及生产环境的日志埋点与稳定性保障。
我们在过去多个企业项目中使用过 ixnode/php-web-crawler 或与其功能相近的方案,如果你在选型或落地过程中遇到问题,例如 版本兼容、二次改造、私有化封装、与内部系统对接、生产 BUG 排查,欢迎联系我们协助评估。
基于 ixnode/php-web-crawler 在你已有业务上做功能扩展、字段裁剪、UI 适配、与内部账号 / 权限 / 日志系统的深度对接。
线上偶发问题、内存泄漏、慢查询、并发异常等排查修复;针对高流量场景做缓存、队列、索引层面的调优。
承接完整的项目从需求 → 设计 → 开发 → 上线 → 长期运维;也可按月提供技术保姆服务。
与 ixnode/php-web-crawler 相关的其它包
同方向 / 同关键字的高下载量 PHP Composer 包推荐,方便对比选型:
Kinikit - PHP Application development framework MVC component
Model of a web-based resource
ext-json wrapper with sane defaults
A package to cast json fields, each sub-keys is castable
Retrieve a WebResourceInterface instance over HTTP
swagger-php - Generate interactive documentation for your RESTful API using phpdoc annotations
统计信息
- 总下载量: 46
- 月度下载量: 0
- 日度下载量: 0
- 收藏数: 2
- 点击次数: 25
- 依赖项目数: 0
- 推荐数: 0
其他信息
- 授权协议: MIT
- 更新时间: 2024-02-24