
SAST应用安全静态分析开发工具代码质量【免费下载链接】semgrepLightweight static analysis for many languages. Find bug variants with patterns that look like source code.项目地址https://gitcode.com/GitHub_Trending/se/semgrep点击查看免费下载本篇技术指南聚焦 Semgrep 仓库中一处小而关键的工程决策在 cli/src/semdep/external/parsy/ 目录下Semgrep 将 Python 解析器组合库 parsy 以 vendor 形式内嵌并为其加入增量行号与列号追踪能力以显著提升依赖锁文件lockfile解析速度。读完本文你将理解offset / line / column三元组索引的改造原理、字符串输入下逐字符增量更新的实现细节、非字符串输入的特殊约定以及行号信息如何被下游 SCA软件组成分析解析器用于错误报告与依赖定位并掌握这一改造与上游合并计划的来龙去脉。一、背景为什么 Semgrep 需要内嵌Vendor一个第三方解析库Semgrep 的依赖解析子系统semdep源码位于 cli/src/semdep/负责扫描各类生态系统的lockfile与manifest文件例如poetry.lock、Pipfile.lock、yarn.lock、go.mod、Gemfile.lock等。从仓库的目录结构可以看到cli/src/semdep/parsers/ 下按生态拆分了大量解析器pipfile.py、poetry.py、yarn.py、go_mod.py、gradle.py、mix.py、swiftpm.py、gem.py、pom_tree.py、requirements.py等它们共同使用 cli/src/semdep/parsers/util.py 提供的通用工具函数。这些解析器并不依赖项目自研的解析框架而是构建在parsy——一个经典的 Python 解析器组合子parser combinator库之上。parsy 的核心思想是通过组合小解析器来构造复杂解析器仓库中的 cli/src/semdep/external/parsy/README.rst 明确写道Parsy 是一个组合小解析器成为复杂、更大解析器的文本解析库属于 LL(infinity) 文法的单子式解析器组合子库。VENDOR_README.md即本次讨论的关联文档点明了内嵌的根本动机Parsy is vendored in order to add incremental line and column tracking functionality, which speeds up our lockfile parsers significantly.即内嵌 parsy 的目的是添加增量式行号与列号追踪功能从而显著加速 Semgrep 的锁文件解析器。由于这一改动尚未被上游接受Semgrep 选择将带补丁的 parsy 直接放入自己的源码树即 vendor而不是通过 PyPI 依赖原版。类似的 vendoring 策略在仓库中还有先例例如 cli/src/semdep/external/packaging/ 同样内嵌了packaging库含specifiers.py、version.py、tags.py等模块。从源码结构看cli/src/semdep/external/parsy/ 一共包含 5 个文件文件作用VENDOR_README.md说明 vendoring 原因、改动设计与上游合并计划__init__.py改造后的 parsy 核心实现解析器、组合子、Position 等__init__.pyiSemgrep 自写的类型桩type stubs约束解析器输入类型version.py版本号当前为2.0cli/src/semdep/external/parsy/version.pyREADME.rst上游 parsy 的原始项目说明文档二、核心改造从单一整数索引到offset / line / column 三元组原版 parsy 在解析字符串时用一个整数索引integer index记录流中已消费了多少个元素。Semgrep 的定制版本将其替换为三个整数组成的索引三元组分别表示 offset偏移、line行号、column列号。VENDOR_README.md 对此给出了精确的定义The edits replace the integer index into a stream with a triple of integers representing offset, line and column information.offset is what index originally was: how many individual elements of the stream have been consumed.即offset就是原索引的含义流的单个元素被消费的个数line与column则是新增的行列信息。这一改造在 cli/src/semdep/external/parsy/init.py 中落地为不可变数据类dataclass(frozenTrue) class Position: offset: int line: int column: intPosition贯穿整个解析流程Parser的包装函数签名变为接收(stream, index: Position)Result中代表解析位置、错误最远位置的字段也从整数升级为Positiondataclass(frozenTrue) class Result: status: bool index: Position value: Any furthest: Position expected: FrozenSet[str] staticmethod def success(index, value): return Result(True, index, value, Position(-1, -1, -1), frozenset()) staticmethod def failure(index, expected): return Result(False, Position(-1, -1, -1), None, index, frozenset([expected]))这里可以观察到两个细节成功结果用Position(-1, -1, -1)占位furthest失败位置字段失败结果则用Position(-1, -1, -1)占位index而把真正的失败位置放进furthest。Result.aggregate负责在alt选择等组合子中合并多个候选解析器的失败信息保留最远失败位置并合并expected集合这正是 parsy 能输出expected ... at line:column式错误信息的基础。对应地ParseError的line_info()方法会按输入类型决定展示方式def line_info(self): if isinstance(self.stream, str): return f{self.index.line}:{self.index.column} else: return str(self.index.offset)即字符串输入报错时显示line:column否则退化为仅显示offset。三、增量更新机制make_index_update 与逐字符行进三元组索引只是数据结构层面的一半另一半在于如何高效地更新它。VENDOR_README.md 描述了更新策略When parser input is a string line and column are updated incrementally with each character of the input string.即在字符串输入下line与column会随输入串的每一个字符被增量地更新。核心实现是make_index_update工厂函数cli/src/semdep/external/parsy/init.pydef make_index_update(consumed: str) - Callable[[Position], Position]: slen len(consumed) line_count consumed.count(\n) last_nl consumed.rfind(\n) return lambda index: Position( offsetindex.offset slen, lineindex.line line_count, columnslen - (last_nl 1) if last_nl 0 else index.column slen, )其数学含义非常清晰offset直接累加已消费字符串的长度slenline累加已消费字符串中换行符\n的数量line_countcolumn分两种情况——若已消费串中包含换行last_nl 0则列号重置为最后一个换行之后剩余的字符数slen - (last_nl 1)若不包含换行则列号直接累加slen。这样每次消费一段字符都只需做常数级的字符串统计count/rfind无需回溯整个已解析前缀重新计算行列。这也是incremental增量式命名的由来与每次需要行列号时都重新扫描输入串的做法相比增量更新将锁文件解析的行列计算开销摊薄到每次消费中从而显著提升解析性能——这正是 VENDOR_README.md 强调的加速效果来源。该工厂函数被三类基础解析器复用1.string解析器——精确匹配一段字符串后用index_update(index)更新位置def string(expected_string: str, transform: Callable[[str], str] noop) - Parser: slen len(expected_string) transformed_s transform(expected_string) index_update make_index_update(expected_string) Parser def string_parser(stream, index): if transform(stream[index.offset : index.offset slen]) transformed_s: return Result.success(index_update(index), expected_string) else: return Result.failure(index, expected_string) return string_parser2.regex解析器——正则匹配成功后对实际匹配到的子串做增量更新def regex_parser(stream, index): match exp.match(stream, index.offset) if match: index ( make_index_update(stream[match.start() : match.end()])(index) if isinstance(stream, str) else Position(match.end(), -1, -1) ) return Result.success(index, match.group(*group))注意正则匹配的长度在匹配前是未知的因此这里用stream[match.start():match.end()]取出真实匹配串再交给make_index_update计算行列增量。3.test_item/test_char解析器——逐字符测试时按单字符更新if isinstance(stream, str): index make_index_update(item)(index) else: index Position(index.offset 1, index.line, index.column)由于make_index_update只处理字符串非字符串路径需要单独分支这正好引出下一节的约定。四、非字符串输入的特殊约定line / column 恒为 -1VENDOR_README.md 还规定了一条重要约定When parser input is not a string (which can never happen with our type stubs) line and column are both set to -1 and ignored.即当解析器的输入不是字符串例如字节流或 token 列表时line与column一律设置为-1并被忽略。文档特别补充说明在 Semgrep 的类型桩约束下这种情况实际上永远不会发生——cli/src/semdep/external/parsy/init.pyi 中的parse、parse_partial等签名都把stream限定为str从静态类型层面保证了调用方只传入字符串。这个约定在parse_partial的入口处落地def parse_partial(self, stream: str | bytes | list): result self( stream, Position(0, 0, 0) if isinstance(stream, str) else Position(0, -1, -1), ) ...字符串输入起点为Position(0, 0, 0)offset、line、column 均从 0 开始非字符串输入起点为Position(0, -1, -1)即 offset 正常累加而 line / column 保持-1表示不可用。前文展示的regex_parser与test_item_parser中非字符串分支分别构造Position(match.end(), -1, -1)与Position(index.offset 1, index.line, index.column)正是为了维持这一约定。而ParseError.line_info()中isinstance(self.stream, str)的分支判断则确保错误信息对非字符串输入回退为纯 offset 展示。五、行号信息的消费端错误报告与依赖定位改造的最终目的是让下游解析器在产出依赖dependency和报错时都能拿到精确的行号。这一能力在 cli/src/semdep/parsers/util.py 中被系统性消费。5.1 零基到一基的转换line_numberparsy 的行列是**零基zero-indexed**的而编辑器和开发者习惯一基1-indexed编号。util.py 通过line_info组合子做了显式转换# parsy line and column numbers are zero indexed, but editors are generally 1 indexed # so we add one to the line number account for this, and discard the column number line_number line_info.map(lambda t: t[0] 1)其中line_info是 vendored parsy 提供的原语解析器返回当前(line, column)二元组见 cli/src/semdep/external/parsy/init.pyline_info Parser(lambda _, index: Result.success(index, (index.line, index.column)))5.2 mark_line为每条解析结果标注行号mark_line在运行目标解析器p之前先取当前行号产出一对(line, result)def mark_line(p: Parser[A]) - Parser[tuple[int, A]]: Returns a parser which gets the current line number, runs [p] and then produces a pair of the line number and the result of [p] return line_number.bind(lambda line: p.bind(lambda x: success((line, x))))它是解析器文档与锁文件条目行号信息的通用来源。以 cli/src/semdep/parsers/poetry.py 为例poetry_dep用mark_line包裹每个[[package]]块的键值对解析并把行号存入ValueLineWrapperpoetry_dep mark_line( string([[package]]\n) key_value_list.map( lambda x: { key_val[0]: ValueLineWrapper(line_number, key_val[1]) for line_number, key_val in x } ) )最终parse_poetry将dep[name].line_number写入FoundDependency.line_number字段把依赖出现在 poetry.lock 的哪一行带进输出output.append( FoundDependency( packagedep[name].value.lower(), versiondep[version].value, ... line_numberdep[name].line_number, lockfile_pathFpath(str(lockfile_path)), manifest_pathFpath(str(manifest_path)) if manifest_path else None, ... ) )cli/src/semdep/parsers/pipfile.py 也采用同一模式用mark_line(key_value).sep_by(new_lines)解析Pipfile清单的键值对并用key_value_list标注行号而对Pipfile.lockJSON 格式则复用 util.py 中基于mark()的 JSON 解析器直接将dep_json.line_number写入FoundDependency。5.3 错误报告ParseError 到 DependencyParserError在 util.py 的parse_dependency_file中捕获ParseError后会把零基的line / column转为一基并连同出错行原文一并封装进DependencyParserErrorexcept ParseError as e: # These are zero indexed but most editors are one indexed line, col e.index.line, e.index.column text_lines text.splitlines() ( [trailing newline] if text.endswith(\n) else [] ) error_str parse_error_to_str(e) if line len(text_lines): offending_line text_lines[line] return DependencyParserError( out.Fpath(str(file_to_parse.path)), file_to_parse.parser_name, error_str, line 1, col 1, offending_line, )这段代码完整利用了改造带来的行列信息不仅能告诉用户哪一行哪一列出错还能直接把出错行原文摘出来放进错误对象极大提升锁文件解析失败时的可诊断性。若行列信息缺失例如解析器在文件开头就失败也会落入专门的分支给出针对性提示。5.4 带行号的 JSON 解析器值得一提的还有 util.py 中从 parsy 官方示例改编而来的 JSON 解析器json_doc。它用mark()组合子包裹每个 JSON 值在解析值前后分别取得(line, column)位置从而让Pipfile.lock等 JSON 锁文件中的每个依赖都能带上行号become( json_value, alt( quoted_str.mark().map(JSON.make), number.mark().map(JSON.make), json_object.mark().map(JSON.make), array.mark().map(JSON.make), true.mark().map(JSON.make), false.mark().map(JSON.make), null.mark().map(JSON.make), ), ) json_doc whitespace json_valueJSON.make取marked[0][0] 1即起始行号 1存入JSON.line_number之后pipfile.py将其透传到FoundDependency.line_number。六、类型桩如何在不改运行时的前提下获得静态类型安全VENDOR_README.md 提到which can never happen with our type stubs指的就是 cli/src/semdep/external/parsy/init.pyi。这份类型桩由 Semgrep 自行编写核心做法是把Parser声明为泛型类Parser(Generic[T])并在parse、parse_partial、__call__等签名中把输入流限定为str从而在 mypy 静态检查层面杜绝非字符串输入路径。有趣的是这种泛型桩与运行时实现之间需要一个小技巧。util.py 开头的模块注释解释得很清楚In Python, type annotations do not do anything, but they are still expressions that get evaluated. The runtime class Parser, as implemented by parsy, takes no parameters, but our type stubs for parsy give this class a generic type variable parameter, so we can enforce staticly that parsers are combined in sensible ways. As a result, the expression Parser[int] is perfectly fine for Mypy, but causes a runtime error. Thankfully, Parser[int] is a perfectly acceptable type annotation for Mypy, and evaluates immediately to string, causing no runtime errors.即运行时Parser不接受类型参数而桩中它是泛型因此源码里不能写Parser[int]这种会被求值、从而触发运行时错误的表达式解决办法是配合from __future__ import annotations使用字符串形式的注解如Parser[str]、Parser[A]让注解在运行时立即求值为字符串mypy 却仍能正确解析其类型含义。这也是 cli/src/semdep/parsers/ 中几乎所有解析器类型注解都写成字符串的原因。类型桩还额外提供了line_info: Parser[Pos]、index: Parser[int]等组合子的签名与运行时实现一一对应。七、与上游的关系向后不兼容的改动与合并计划VENDOR_README.md 的最后一段交代了这一改动与上游 parsy 的关系这是理解为什么选择 vendoring 而非提 PR 等合入的关键Matthew is planning to merge this change upstream into parsy, but its currently backwards incompatible, so were vendoring until thats solved, the change is merged, and parsy is released.也就是说Semgrep 团队成员 Matthew 计划把这一改动合入上游 parsy但改动目前是向后不兼容的把公开的索引语义从整数改为三元组Position会破坏依赖原 API 的用户代码因此在兼容性问题解决、改动合入并发布新版 parsy 之前Semgrep 选择持续 vendoring这份定制实现。配套的版本信息可以在 cli/src/semdep/external/parsy/version.py 中看到__version__ 2.0。这一版本号既标识了 vendored 实现所对应的 parsy 基线版本也方便后续跟踪替换为上游正式版本时的 diff。对想要替换回上游依赖的开发者而言对比__init__.py中Position与make_index_update的实现即可精确识别出 Semgrep 的增量行列补丁范围。八、总结与延伸阅读Semgrep 对 parsy 的 vendoring 是一次典型的带着私有补丁的第三方依赖管理实践以最小侵入仅替换索引表示 增量行列更新换取锁文件解析的显著加速同时用类型桩守住类型安全底线用明确的文档记录改动语义与上游合并计划为将来回迁上游版本铺好了路。其增量更新思路用count(\n)/rfind(\n)做常数级行列推进也可以作为其他文本解析器性能优化的通用参考。若想继续深入可以在仓库中按以下路径追踪完整链路改造文档cli/src/semdep/external/parsy/VENDOR_README.md改造实现Position、Result、make_index_update、string、regex、test_item、line_info等均位于 cli/src/semdep/external/parsy/init.py类型桩cli/src/semdep/external/parsy/init.pyi行号消费工具line_number、mark_line、json_doc、parse_dependency_file的错误处理cli/src/semdep/parsers/util.py典型消费方cli/src/semdep/parsers/poetry.py、cli/src/semdep/parsers/pipfile.py其余基于 parsy 的解析器cli/src/semdep/parsers/yarn.py、cli/src/semdep/parsers/go_mod.py、cli/src/semdep/parsers/gradle.py、cli/src/semdep/parsers/mix.py、cli/src/semdep/parsers/swiftpm.py、cli/src/semdep/parsers/gem.py、cli/src/semdep/parsers/pom_tree.py、cli/src/semdep/parsers/requirements.py赞分享SAST应用安全静态分析开发工具代码质量【免费下载链接】semgrepLightweight static analysis for many languages. Find bug variants with patterns that look like source code.项目地址https://gitcode.com/GitHub_Trending/se/semgrep点击查看免费下载相关推荐Flutter News Toolkit数据管理缓存策略与离线功能实现教程Flutter News Toolkit数据管理缓存策略与离线功能实现教程 Flutter News Toolkit是由Google和Very Good Ve3 步跑通 Hermes Agent自进化 AI 代理安装配置完全指南3 步跑通 Hermes Agent自进化 AI 代理安装配置完全指南 本文带你装好 Hermes Agent——一个能从经验里沉淀技能、越用越懂你的 AIAI Agent人工智能AI 应用工具调用Agent 记忆交互助手RAG任务调度MCP 服务Puppeteer HTTPRequest.redirectChain() 深入解析如何追踪与检测页面重定向链路Puppeteer HTTPRequest.redirectChain 深入解析如何追踪与检测页面重定向链路 导读 重定向Redirect是 Web 中最浏览器控制测试网页爬虫开发工具上一篇Zotero重复文献终极清理指南5分钟掌握智能合并插件的完整使用技巧下一篇Zotero重复文献智能合并插件三步实现文献库高效整理创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考