interitty/tokenizer
Composer 安装命令:
composer require interitty/tokenizer
包简介
Use regular expressions to split a given string into tokens.
关键字:
README 文档
README
Use regular expressions to split a given string into tokens.
Requirements
- PHP >= 8.5
Installation
The best way to install interitty/tokenizer is using Composer:
composer require interitty/tokenizer
Tokenizer usage
The tokenization process needs the definition of a map (from token regexes to token classes) and string to be tokenized.
A simple tokenizer that separates strings into numbers, whitespaces, and letters can look like the following code.
$tokenizer = new Tokenizer('say 123');
$tokenizer->map = [
'number' => '~^\d+~',
'whitespace' => '~^\s+~',
'string' => '~^\w+~'
];
Processing the tokens
Tokens can be accessed by iterating thru the next and current methods until the TOKEN_END appears.
$tokens = [];
do {
$token = $tokenizer->next();
$tokens[] = $token;
assert($token === $tokenizer->current());
} while ($token->getType() !== Token::TOKEN_END);
The resulting array of $tokens would look like the following.
[
new Token('string', 'say', 1, 1),
new Token('whitespace', ' ', 1, 4),
new Token('number', '123', 1, 5),
]
Skipping unnecessary tokens
In some cases, it may be useful to automatically skip some tokens and move on to others.
Because of that, there are addSkippedTokenType and setSkippedTokenTypes methods.
The TOKEN_END token can't be skipped.
$tokenizer->addSkippedTokenType('whitespace');
$string = '';
do {
$token = $tokenizer->next();
$string .= $token->getValue();
} while ($token->getType() !== Token::TOKEN_END);
assert('say123' === $string);
Expecting tokens
The tokenizer includes a helper to expect the correct token type and value. This can simplify and unify the checking process.
$tokenizer = new Tokenizer('{some coed}');
$tokenizer->map = [
'brackets' => '~^[{}]~',
'code' => '~^[^{}]+~',
];
$tokenizer->expect($tokenizer->next(), 'brackets', '{');
$tokenizer->expect($tokenizer->next(), 'code');
$code = $tokenizer->current()->getValue();
$tokenizer->expect($tokenizer->next(), 'brackets', '}');
BaseTokenizerParser usage
The possible way for using a Tokenizer is in the BaseTokenizerParser which provides the functionality of parsing
the given string into a stream of tokens. It can be useful for validating that a given string is compatible with the
expected grammar and for parsing him into a structured array.
This functionality is used in the interitty/pacc.
BaseParser usage
In the case where it can be needed to work with own implementation of Tokenizer, there is a BaseParser abstract
class that allows implementing own logic of work with current and next Token and own mechanism of work with the
tokenType and tokenLexeme.
interitty/tokenizer 适用场景与选型建议
interitty/tokenizer 是一款 基于 PHP 开发的 Composer 扩展包,目前已累计 52 次下载、GitHub Stars 达 1, 最近一次更新时间为 2022 年 08 月 26 日, 在 PHP 生态内属于活跃度较高的组件。
它主要适用于以下技术方向: 「php」 「tokenizer」 「yacc」 「interitty」 「pacc」 等业务场景。在实际项目中,围绕这些方向常见需要落地的问题包括:接口对接、性能调优、并发安全、与既有框架(Laravel / ThinkPHP / Yii / Webman 等)的兼容适配,以及生产环境的日志埋点与稳定性保障。
我们在过去多个企业项目中使用过 interitty/tokenizer 或与其功能相近的方案,如果你在选型或落地过程中遇到问题,例如 版本兼容、二次改造、私有化封装、与内部系统对接、生产 BUG 排查,欢迎联系我们协助评估。
基于 interitty/tokenizer 在你已有业务上做功能扩展、字段裁剪、UI 适配、与内部账号 / 权限 / 日志系统的深度对接。
线上偶发问题、内存泄漏、慢查询、并发异常等排查修复;针对高流量场景做缓存、队列、索引层面的调优。
承接完整的项目从需求 → 设计 → 开发 → 上线 → 长期运维;也可按月提供技术保姆服务。
与 interitty/tokenizer 相关的其它包
同方向 / 同关键字的高下载量 PHP Composer 包推荐,方便对比选型:
Library emulating the PHP internal reflection using just the tokenized source code.
A program for performing lexical analysis, written in PHP
A simple library for make Lexer and Parsers to build a language
Wrapper around PHP's tokenizer extension.
Html5 stream tokenizer/reader (not using libxml)
Rudl UDP cluster logging library (PSR4 compliant)
统计信息
- 总下载量: 52
- 月度下载量: 0
- 日度下载量: 0
- 收藏数: 2
- 点击次数: 27
- 依赖项目数: 1
- 推荐数: 0
其他信息
- 授权协议: MIT
- 更新时间: 2022-08-26