多彩编程 多彩编程MZPH · CODE BLOG
ARTICLE DETAIL

文章详情

深耕前端与后端开发技术的一线实战笔记与踩坑复盘。

深入解析 Apache Zeppelin Python 与 IPython 解释器:进程架构、Py4J 桥接与 gRPC 内核通信原理

深入解析 Apache Zeppelin Python 与 IPython 解释器:进程架构、Py4J 桥接与 gRPC 内核通信原理 后端前端大数据数据分析【免费下载链接】zeppelinWeb-based notebook that enables>项目地址https://gitcode.com/gh_mirrors/zeppelin2/zeppelin点击查看免费下载导读本文以仓库 python/README.md 为骨架完整解析 Apache Zeppelin 中 Python 解释器的两代实现——基于ProcessBuilder Py4J 的原生 Python 解释器以及基于jupyter_client gRPC 的 IPython 解释器。你将掌握解释器进程如何被拉起与通信、段落执行/中断/结果回传的底层协议、matplotlib 内联渲染与%python.sql的实现机制、全部可配置参数及其默认值以及如何运行依赖真实 Python 环境的完整测试套件。一、Python 解释器在 Zeppelin 中的定位Apache Zeppelin 的 notebook 支持以 SQL、Scala 等多种语言编写数据驱动、交互式的分析段落Python 是其最常用的解释器之一。按 python/README.md 的说明该模块提供两大能力Python interpreter经典实现通过系统 Python 进程执行代码支持 Py4J 动态表单、matplotlib 内联绘图、%python.sqlPandas DataFrame 上的 SQLIPython interpreter将整个执行工作委托给 IPython kernel通过 gRPC 通信尽量让所有 IPython 生态特性如自动补全、富输出在 Zeppelin 中原样可用。整个python模块位于仓库 python/ 目录核心 Java 类集中在 python/src/main/java/org/apache/zeppelin/python/PythonInterpreter、IPythonInterpreter、PythonInterpreterPandasSql、PythonZeppelinContext等Python 侧引导代码在 python/src/main/resources/python/ 与 python/src/main/resources/grpc/python/。二、原生 Python 解释器的总体架构README 原文Current interpreter implementation spawns new system python process throughProcessBuilderand re-directs its stdin\strout to Zeppelin当前实现通过ProcessBuilder派生一个系统 Python 进程并将其 stdin/stdout 重定向到 Zeppelin。整体数据流如下Zeppelin 解释器进程JVM在open()时启动一个常驻sleeping的 Python 进程用户的段落代码由 JVM 写入 Python 进程的 stdinPython 执行结果、stdout 输出被重定向回 Zeppelin并实时流转到 notebook 段落输出区JVM 与 Python 进程之间同时建立一条Py4J 双向通道Python 侧通过 JavaGateway 反向调用 JVM 对象如动态表单、上下文JVM 侧通过 gateway 的entry_point暴露解释器实例。2.1 进程拉起细节-i与-u选项按 README 的 Technical overview 描述When interpreter is starting it launches a python process inside a Java ProcessBuilder. Python is started with-i(interactive mode) and-u(unbuffered stdin, stdout and stderr) options. Thus the interpreter has a sleeping python process.启动时以-i交互模式与-u无缓冲模式启动 Python因此解释器持有一个休眠中的 Python 进程。-u保证 stdout/stderr 不缓冲使 JVM 侧能实时读取输出-i让进程进入交互 REPL 状态等待指令避免每条语句单独拉起进程、从而在段落间保留变量与状态。在 PythonInterpreter.java 的createGatewayServerAndStartScript()中可以看到具体实现将资源脚本python/zeppelin_python.py复制为临时文件构造器中创建于/tmp见 PythonInterpreter.java通过findRandomOpenPortOnAllLocalInterfaces()随机选择一个空闲端口启动 Py4JGatewayServer监听0.0.0.0用CommandLine.parse(getPythonCommand())构造命令行并追加脚本路径、端口、本机 IP 三个参数交给 Apache Commons Exec 的DefaultExecutor异步执行设置PYTHONPATH环境变量把py4j-0.9.2/src与interpreter/lib/python内含 matplotlib 内联后端backend_zinline.py、mpl_config.py注入 Python 进程见 PythonInterpreter.java 与 interpreter/lib/python/使用ExecuteWatchdog(INFINITE_TIMEOUT)守护进程确保 Python 进程在整个解释器生命周期内持续存活。2.2 请求/响应协议从 flush marker 到 Py4J 回调README 记载的经典协议为Interpreter sends command to python with a JavaoutputStreamWriterand read from anInputStreamReader. To know when stop reading stdout, interpreter sendsprint *!?flush reader!?*after each command and reads stdout until he receives back the*!?flush reader!?*.JVM 用outputStreamWriter发送命令、用InputStreamReader读取输出每条命令后发送print *!?flush reader!?*直到读回该标记才停止读取。需要说明的是在当前仓库快照的实现中段落请求与结果回传改由 Py4J 回调完成Python 侧主循环通过intp.getStatements()阻塞等待 JVM 下发代码执行完毕后调用intp.setStatementsFinished(out, error)回传结果与错误状态见 python/src/main/resources/python/zeppelin_python.pyJVM 侧对应方法在 PythonInterpreter.java。整个握手围绕statementSetNotifier、statementFinishedNotifier、pythonScriptInitializeNotifier三个同步对象完成例如interpret()在等待 Python 完成初始化时最多阻塞MAX_TIMEOUT_SEC 10秒超时则返回 python is not responding见 PythonInterpreter.java。可以推断 flush marker 协议对应更早期的实现README 描述的是该机制的设计意图——即如何界定一段 stdout 的读取边界。2.3 Bootstrap 引导与z上下文对象README 指出When interpreter is starting, it sends some Python code (bootstrap.py and bootstrap_input.py) to initialize default behavior and functions (help(), z.input()...). bootstrap_input.py is sent only if py4j library is detected inside Python process.启动时发送 bootstrap 代码初始化默认行为与函数bootstrap_input.py仅在 Python 进程内检测到 py4j 库时才发送。在当前快照中该职责由 zeppelin_python.py 承担通过GatewayClient(addresshost, portint(sys.argv[1]))连接 JVM 侧 gatewaygateway.entry_point即解释器实例定义PyZeppelinContext别名z/__zeppelin__向 Python 用户空间暴露z.input(name, defaultValue)、z.textbox、z.noteTextbox动态表单输入z.select、z.noteSelect、z.checkbox、z.noteCheckbox选项/复选表单选项通过gateway.jvm构造 Java 侧的OptionInput.ParamOption数组z.show(df)以 Zeppelin%table显示系统渲染 DataFrame默认最多输出max_result1000行z.show(plt)以 base64 PNG/SVG 内联渲染 matplotlib 图z.registerHook / unregisterHook / registerNoteHook注册执行钩子如post_exec。语句按 AST 分别以exec与single模式编译执行使最后一条表达式的求值结果能打印到 stdout模拟 REPL 行为见 zeppelin_python.py段落的用户变量保存在独立命名空间_zcUserQueryNameSpace与解释器内部命名隔离。2.4 段落中断SIGINT 信号README 原文JavaBuilder cant send SIGINT signal to interrupt paragraph execution. Therefore interpreter directly send akill SIGINT PIDto python process to interrupt execution. Python process catch SIGINT signal with some code defined in bootstrap.py.JVM 无法直接向 Python 进程发送 SIGINT因此解释器直接执行kill SIGINT PIDPython 进程在 bootstrap 中注册 SIGINT 处理器。源码印证JVM 侧interrupt()在pythonPid -1时执行Runtime.getRuntime().exec(kill -SIGINT pythonPid)对非 UNIX/Linux 系统则回退为直接关闭解释器见 PythonInterpreter.java。pythonPid由 Python 进程启动时通过intp.onPythonScriptInitialized(os.getpid())上报Python 侧在引导代码中注册signal.signal(signal.SIGINT, handler_stop_signals)收到信号后抛出带信号信息的异常结束当前段落见 zeppelin_python.py。2.5 matplotlib 内联渲染README 原文Matplotlib figures are displayed inline with the notebook automatically using a built-in backend for zeppelin in conjunction with a post-execute hook.使用 Zeppelin 内置后端 post-execute 钩子自动将 matplotlib 图形内联显示在 notebook 中。实现要点JVM 在open()时注册POST_EXEC_DEV钩子__zeppelin__._displayhook()见 PythonInterpreter.javaPython 侧_setup_matplotlib()优先将 matplotlib 切到module://backend_zinline位于 interpreter/lib/python/backend_zinline.py并调用 interpreter/lib/python/mpl_config.py 配置width600, height400, dpi72, fontsize10, interactiveTrue, formatpng若后端缺失则回退到 Agg见 zeppelin_python.pyshow_matplotlib()将图形编码为data:image/png;base64,...的imgHTML以%html输出。2.6%python.sql对 Pandas DataFrame 执行 SQLREADME 原文%python.sqlsupport for Pandas DataFrames is optional and provided using pandasql if user have one installed.%python.sql对 Pandas DataFrame 的 SQL 支持是可选的依赖用户安装 pandasql。SQL 引导脚本 python/src/main/resources/python/bootstrap_sql.py 定义pysqldf lambda q: sqldf(q, globals())若未安装 pandas/pandasql则给出友好提示而非报错中断PythonInterpreterPandasSql.open()通过bootStrapInterpreter(/python/bootstrap_sql.py)注入该函数interpret()将用户的 SQL 包装为__zeppelin__.show(pysqldf(...))委托给PythonInterpreter执行见 PythonInterpreterPandasSql.java。其注释明确目的是复刻%spark.sql对 Spark DataFrame 的体验。三、解释器配置参数速查依据 python/src/main/resources/interpreter-setting.json%python组共注册五个解释器python、ipython、sql、conda、docker其中可配置参数如下解释器参数默认值说明pythonzeppelin.pythonpythonPython 可执行文件路径默认假定python在$PATH中pythonzeppelin.python.maxResult1000DataFrame 表格显示的最大行数pythonzeppelin.python.useIPythontrue当 IPython 可用时是否优先使用 IPython 解释器ipythonzeppelin.ipython.launch.timeout30000IPython kernel 启动超时毫秒ipythonzeppelin.ipython.grpc.message_size3355443232MgRPC 消息大小上限字节代码层面对应关系zeppelin.python.maxResult在 PythonInterpreter.java 解析并传入PythonZeppelinContextzeppelin.python在getPythonBindPath()中解析缺省回退pythonPythonInterpreter.javazeppelin.ipython.launch.timeout与zeppelin.ipython.grpc.message_size均在 IPythonInterpreter.java 读取。四、开发前提与测试4.1 Dev prerequisites按 python/README.md每台机器需安装Python 2 或 3且装有py4j0.9.2与matplotlib1.31 或更高单元测试只校验解释器逻辑不会启动真实 Python 进程——Python 进程被一个把输入原样输出的 mock 类替代写在bootstrap.py/bootstrap_input.py当前快照对应 zeppelin_python.py 与 bootstrap_sql.py中的代码必须同时兼容 Python 2 与 3Python 代码遵循PEP8规范。4.2 运行完整测试套件默认构建会跳过依赖真实 Python 环境与外部库Pandas、Pandasql 等的用例要运行包括这些用例在内的全部测试执行 README 给出的命令mvn -Dpython.test.exclude test -pl python -am其中-Dpython.test.exclude对应 python/pom.xml 中 surefire 插件的excludes配置默认排除值使慢速用例跳过-pl python -am表示仅构建 python 模块及其依赖模块。相关测试类位于 python/src/test/java/org/apache/zeppelin/python/如PythonInterpreterTest、PythonInterpreterMatplotlibTest、PythonInterpreterPandasSqlTest、IPythonInterpreterTest等。五、IPython 解释器架构与要求5.1 依赖要求按 python/README.mdIPython 解释器正常工作需要安装以下 Python 包jupyter 5.xIPythonipykernelgrpcio若已安装 Anaconda则只需额外安装grpc其余包自带。源码中的前置检查更严格checkIPythonPrerequisite()会执行pip freeze并依次确认jupyter-client、ipykernel、ipython、grpcio、protobuf均已安装任一缺失即返回对应错误信息见 IPythonInterpreter.java。5.2 架构jupyter_client gRPCREADME 原文Current interpreter delegate the whole work to ipython kernel viajupyter_client. Zeppelin would launch a python process which host the ipython kernel. Zeppelin interpreter process will communicate with the python process viagrpc. Ideally every feature works in IPython should work in Zeppelin as well.解释器通过jupyter_client把全部工作委托给 IPython kernelZeppelin 启动一个承载 IPython kernel 的 Python 进程并通过 gRPC 与之通信。理论上 IPython 中可用的每个特性都应在 Zeppelin 中可用。具体链路依据 IPythonInterpreter.java 与 python/src/main/resources/grpc/python/open()随机分配两个端口ipythonPortgRPC 服务与jvmGatewayPortPy4J gateway将ipython_server.py、ipython_pb2.py、ipython_pb2_grpc.py复制到临时目录用python ipython_server.py port拉起承载 IPython kernel 的进程见 IPythonInterpreter.java启动 JVM 侧 Py4JGatewayServer并通过 gRPC 向 kernel 注入 grpc/python/zeppelin_python.py占位符${JVM_GATEWAY_PORT}被替换为实际端口打通 Python ↔ JVM 的动态表单通道见 IPythonInterpreter.java启动后轮询status直到 kernel 处于RUNNING超过zeppelin.ipython.launch.timeout默认 30 秒则报错退出。5.3 gRPC 协议定义通信契约定义在 python/src/main/proto/ipython.proto由 protobuf-maven-plugin 编译生成 Java 存根依赖grpc 1.15.0见 python/pom.xml。IPython服务共五个 RPCexecute(ExecuteRequest) returns (stream ExecuteResponse)执行代码以流式返回输出支持 TEXT / IMAGE 两种输出类型complete(CompletionRequest) returns (CompletionResponse)代码补全cancel(CancelRequest) returns (CancelResponse)取消正在运行的语句status(StatusRequest) returns (StatusResponse)查询 kernel 状态STARTING/RUNNINGstop(StopRequest) returns (StopResponse)关闭 kernel。段落执行时IPythonInterpreter.interpret()调用ipythonClient.stream_execute(...)将执行输出经InterpreterOutputStream实时写入段落最终根据ExecuteStatusSUCCESS/ERROR映射为InterpreterResult见 IPythonInterpreter.java。补全结果在completion()中被解析为 Zeppelin 的InterpreterCompletion列表因此支持TAB键补全见 interpreter-setting.json 中 ipython 的completionKey: TAB。5.4 与原生解释器的自动切换PythonInterpreter.open() 展示了两代实现的衔接逻辑若zeppelin.python.useIPythontrue默认且checkIPythonPrerequisite()返回空前置满足则优先打开IPythonInterpreter其后interpret/cancel/progress/completion全部委托给它若 IPython 不可用或打开失败自动回退到原生PythonInterpreter此时才注册 matplotlib 钩子并拉起 gateway Python 进程。该设计让用户在未安装 IPython 生态包的环境下仍能正常使用%python无需手工切换配置。子类还可通过setAdditionalPythonPath/setAdditionalPythonInitFile/setAddBulitinPy4j扩展如 PySpark 场景追加 PYTHONPATH 与初始化代码见 IPythonInterpreter.java。六、小结Apache Zeppelin 的 Python 支持呈现了一条清晰的演进路径经典 Python 解释器进程级集成——ProcessBuilder拉起带-i -u的常驻 PythonPy4J 双向桥接提供动态表单与上下文SIGINT 实现段落中断内置 backend post-exec 钩子实现 matplotlib 内联pandasql 带来%python.sqlIPython 解释器内核级集成——由jupyter_client托管 IPython kernelgRPC 流式执行代码并回传输出保留了补全、取消、状态查询等完整能力并在前置依赖不满足时优雅回退到经典实现。无论是阅读 python/README.md 快速上手还是深入 PythonInterpreter.java、IPythonInterpreter.java、ipython.proto 与 zeppelin_python.py 验证实现细节理解上述两条通信链路与配置参数是排查%python执行异常、定制 Python 环境与扩展解释器能力的起点。赞分享后端前端大数据数据分析【免费下载链接】zeppelinWeb-based notebook that enables>项目地址https://gitcode.com/gh_mirrors/zeppelin2/zeppelin点击查看免费下载相关推荐Apache Zeppelin Python 解释器深度解析进程架构、Py4J 桥接与 IPython 内核集成Apache Zeppelin Python 解释器深度解析进程架构、Py4J 桥接与 IPython 内核集成 Apache Zeppelin 的 Pyth数据分析数据可视化大数据后端前端任务调度如何快速获取网盘直链LinkSwift 网盘直链下载助手完整使用指南如何快速获取网盘直链LinkSwift 网盘直链下载助手完整使用指南 LinkSwift 是一个浏览器用户脚本作用是在百度网盘、阿里云盘等九大网盘的网页端帮数据分析数据可视化大数据后端Apache Zeppelin Python 解释器实战指南%python、IPython 与 Pandas SQL 全解析Apache Zeppelin Python 解释器实战指南%python、IPython 与 Pandas SQL 全解析 导读 本文以 Apache Z数据分析数据可视化大数据后端上一篇commitlint错误排查手册常见问题及解决方案大全下一篇ClickVisual用户指南从安装到高级查询的10个实用技巧创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表