Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions docs/en/abbreviations.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ This page lists high-frequency technical abbreviations used throughout the book

## General Abbreviations

*Table FM-1: General Abbreviations.*
| Abbreviation | Full Name | Description | Main Locations |
| --- | --- | --- | --- |
| A100 | NVIDIA A100 GPU | NVIDIA A100 accelerator | Part 1, Part 10, Part 11 |
Expand Down Expand Up @@ -37,6 +38,7 @@ This page lists high-frequency technical abbreviations used throughout the book

## Data Engineering and Platforms

*Table FM-2: Data Engineering and Platforms.*
| Abbreviation | Full Name | Description | Main Locations |
| --- | --- | --- | --- |
| DataOps | Data Operations | Data operations and data-engineering operations system | Part 2, Part 8, Part 10 |
Expand All @@ -58,6 +60,7 @@ This page lists high-frequency technical abbreviations used throughout the book

## Training, Alignment, and Reasoning

*Table FM-3: Training, Alignment, and Reasoning.*
| Abbreviation | Full Name | Description | Main Locations |
| --- | --- | --- | --- |
| CoT | Chain-of-Thought | Chain-of-thought reasoning | Part 6, Part 10, Part 11 |
Expand All @@ -76,6 +79,7 @@ This page lists high-frequency technical abbreviations used throughout the book

## Multimodality and Vision

*Table FM-4: Multimodality and Vision.*
| Abbreviation | Full Name | Description | Main Locations |
| --- | --- | --- | --- |
| BBox | Bounding Box | Bounding box | Part 3, Part 10, Part 11 |
Expand All @@ -99,6 +103,7 @@ This page lists high-frequency technical abbreviations used throughout the book

## Evaluation, Compliance, and Governance

*Table FM-5: Evaluation, Compliance, and Governance.*
| Abbreviation | Full Name | Description | Main Locations |
| --- | --- | --- | --- |
| AGI-Eval | AGI Evaluation | Evaluation benchmark for general-intelligence capabilities | Part 11 |
Expand Down
8 changes: 4 additions & 4 deletions docs/en/appendix_a_tools_and_frameworks_quick_reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,7 +253,7 @@ The later parts of the book connect naturally to tool choices:
- Chapters 44-45: pre-training and post-training recipes need batch processing, version governance, experiment tracking, and data cards.
- Chapters 46-48: reasoning, multimodal, and generative scenarios need trajectory records, evaluation slices, storage layering, and inference services.
- Chapters 38-43: specialized datasets need fact checking, sample schemas, build pipelines, evaluation protocols, compliance audits, and reproducibility boundaries.
- Appendices A-C translate those capabilities into operational checklists and templates for project managers, teaching assistants, platform teams, and maintainers.
- Appendices A-H translate those capabilities into operational checklists and templates for project managers, teaching assistants, platform teams, and maintainers.

This reminds readers that appendices are not secondary extras. They translate engineering capabilities from the main text into operational language.

Expand Down Expand Up @@ -292,8 +292,8 @@ Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Ra

Pushkarna M, Zaldivar A, Kjartansson O, Cicconi P, Chen V, Efrat A, Zou Y, Mueller J, Taly A, Ehyaei A, Karkkainen K, Marathe A, Han X, Mittal A, Schuster T, Yarmand M, Sohn H, Dwarakanath N C, McCann B (2022) Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. In: Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp 1776-1826. https://doi.org/10.1145/3531146.3533231.

DVC Contributors (2026) Data Version Control Documentation. Available at: https://dvc.org/doc.
DVC Contributors (2026) Data Version Control Documentation. https://dvc.org/doc.

MLflow Authors (2026) MLflow Documentation. Available at: https://mlflow.org/docs/latest/.
MLflow Authors (2026) MLflow Documentation. https://mlflow.org/docs/latest/.

Hugging Face (2026) Hugging Face Datasets Documentation. Available at: https://huggingface.co/docs/datasets.
Hugging Face (2026) Hugging Face Datasets Documentation. https://huggingface.co/docs/datasets.
10 changes: 5 additions & 5 deletions docs/en/appendix_b_compliance_and_release_checklist.md
Original file line number Diff line number Diff line change
Expand Up @@ -311,14 +311,14 @@ Third, long-term risk is reduced not by one approval form but by source records,

## References

National People's Congress of the People's Republic of China (2016) Cybersecurity Law of the People's Republic of China. Available at: https://www.gov.cn/xinwen/2016-11/07/content_5129723.htm.
National People's Congress of the People's Republic of China (2016) Cybersecurity Law of the People's Republic of China. https://www.gov.cn/xinwen/2016-11/07/content_5129723.htm.

National People's Congress of the People's Republic of China (2021a) Data Security Law of the People's Republic of China. Available at: https://www.gov.cn/xinwen/2021-06/11/content_5616919.htm.
National People's Congress of the People's Republic of China (2021a) Data Security Law of the People's Republic of China. https://www.gov.cn/xinwen/2021-06/11/content_5616919.htm.

National People's Congress of the People's Republic of China (2021b) Personal Information Protection Law of the People's Republic of China. Available at: https://www.gov.cn/xinwen/2021-08/20/content_5632486.htm.
National People's Congress of the People's Republic of China (2021b) Personal Information Protection Law of the People's Republic of China. https://www.gov.cn/xinwen/2021-08/20/content_5632486.htm.

National Institute of Standards and Technology (2023) AI Risk Management Framework (AI RMF 1.0). Available at: https://www.nist.gov/itl/ai-risk-management-framework.
National Institute of Standards and Technology (2023) AI Risk Management Framework (AI RMF 1.0). https://www.nist.gov/itl/ai-risk-management-framework.

European Parliament and Council of the European Union (2024) Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Available at: https://eur-lex.europa.eu/eli/reg/2024/1689/oj.
European Parliament and Council of the European Union (2024) Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). https://eur-lex.europa.eu/eli/reg/2024/1689/oj.

Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji I D, Gebru T (2019) Model Cards for Model Reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency, pp 220-229. https://doi.org/10.1145/3287560.3287596.
4 changes: 2 additions & 2 deletions docs/en/appendix_c_cost_estimation_and_resource_templates.md
Original file line number Diff line number Diff line change
Expand Up @@ -319,6 +319,6 @@ Narayanan D, Shoeybi M, Casper J, LeGresley P, Patwary M, Catanzaro B (2021) Eff

Kwon W, Li Z, Zhuang S, Sheng Y, Zheng L, Yu C H, Gonzalez J E, Zhang H, Stoica I (2023) Efficient Memory Management for Large Language Model Serving with PagedAttention. In: Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, pp 611-626. https://doi.org/10.1145/3600006.3613165.

Kubernetes Authors (2026) Kubernetes Documentation. Available at: https://kubernetes.io/docs/.
Kubernetes Authors (2026) Kubernetes Documentation. https://kubernetes.io/docs/.

vLLM Project (2026) vLLM Documentation. Available at: https://docs.vllm.ai/.
vLLM Project (2026) vLLM Documentation. https://docs.vllm.ai/.
2 changes: 1 addition & 1 deletion docs/en/appendix_d_paper_to_implementation_guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,7 +178,7 @@ For a chapter-level reproduction repository, consider this structure:
| `data/` | Data snapshot, index, and version note |
| `src/` | Core implementation |
| `configs/` | Parameters, paths, run configuration |
| `scripts/` | One-command run scripts |
| `scripts/` | Reproducible run scripts |
| `eval/` | Evaluation scripts and slice reports |
| `docs/` | Documentation, boundary notes, FAQ |
| `reports/` | Result charts, acceptance screenshots, postmortems |
Expand Down
24 changes: 12 additions & 12 deletions docs/en/appendix_g_datagallery_note.md
Original file line number Diff line number Diff line change
@@ -1,32 +1,32 @@
# Appendix G: DataGallery Open-source Ecosystem Overview
# Appendix G: DataGallery Open-source Ecosystem and Reproduction Notes

## G.1 Purpose of This Appendix

This appendix explains where DataGallery sits in Project 15 and in the agentic data engineering practices discussed in this book. The public open-source entry for DataGallery is hosted on GitCode at [https://gitcode.com/datagallery](https://gitcode.com/datagallery). This appendix is not an installation manual for DataGallery or DataAgent, nor does it replace the README, example configurations, dependency notes, or release records in the corresponding repositories. When reproducing an experiment or integrating the project into an engineering system, readers should treat the public repository, a specific tag or commit, and the project documentation as the source of truth (DataGallery Contributors 2026a; DataGallery Contributors 2026b).
This appendix explains where DataGallery sits in Project 15 and in the agentic data engineering practices discussed in this book. The public open-source entry for DataGallery is hosted on GitCode at [https://gitcode.com/datagallery](https://gitcode.com/datagallery). This appendix is not an installation manual for DataGallery or DataAgent, nor does it replace the README, example configurations, dependency notes, or release records in the corresponding repositories. When reproducing an experiment or integrating the project into an engineering system, readers should treat the public repository, a specific tag or commit, and the project documentation as the source of truth.

In this book, DataGallery is best understood as a set of open-source engineering entry points around Data + AI practice, rather than as the name of one isolated tool. Project 15 uses DataAgent to build an enterprise semantic BI assistant. Its core question is not "how to call a model to generate SQL," but how to organize business questions, the semantic layer, an NL2SQL sub-agent, tool calls, workspace assets, runtime traces, and service interfaces into a reviewable data engineering system. DataGallery provides the open-source home for this kind of work, while DataAgent is the project entry most directly referenced in this book.
In this book, DataGallery is best understood as a set of open-source engineering entry points around Data + AI practice, rather than as the name of one isolated tool. Project 15 uses DataAgent to build an enterprise semantic BI assistant. Its core question is not "how to call a model to generate SQL," but how to organize business questions, the semantic layer, an NL2SQL sub-agent, tool calls, workspace assets, runtime traces, and service interfaces into a reviewable data engineering system. DataGallery provides the public organization and project entry, while DataAgent is the engineering project most directly referenced in this book.

This appendix therefore focuses on the relationship between DataGallery and the book's data engineering methods: how it helps readers connect the chapters on agents, tool use, semantic layers, DataOps, reproduction, and governance to runnable, auditable, and iterative open-source projects. For concrete APIs, startup commands, configuration fields, and dependency versions, this appendix gives usage principles rather than a step-by-step tutorial.
This appendix therefore focuses on the relationship between DataGallery and the book's data engineering methods: how the chapters on agents, tool use, semantic layers, DataOps, reproduction, and governance can be connected to runnable, auditable, and iterative open-source projects. For concrete APIs, startup commands, configuration fields, and dependency versions, this appendix gives usage principles rather than a step-by-step tutorial.

## G.2 Relationship to Project 15

Project 15 uses DataAgent as the practical object for an enterprise semantic BI assistant because it places agent orchestration, semantic-layer enhancement, NL2SQL, tool use, workspace assets, and service interfaces in one engineering chain. DataGallery provides the higher-level open-source organization entry, so DataAgent can be understood not as an isolated repository, but as part of an ecosystem for data agents, data applications, and reproducible data engineering.

This distinction matters. DataAgent is the concrete project: readers need to inspect its YAML configuration, Semantic Service, NL2SQL sub-agent, execution tools, workspace, and A2A interface. DataGallery is the ecosystem entry: readers should use it to confirm project ownership, public repositories, license status, maintenance state, related projects, and future migration clues. For a book chapter, the main text should explain the project chain; an appendix is the right place to explain the open-source ecosystem and reproduction boundary.
This boundary should remain explicit. DataAgent is the concrete project: reproduction requires checking its YAML configuration, Semantic Service, NL2SQL sub-agent, execution tools, workspace, and A2A interface. DataGallery is the ecosystem entry: readers can use it to confirm project ownership, public repositories, license status, maintenance state, related projects, and future migration clues. For a book chapter, the main text should explain the project chain; an appendix is the right place to explain the open-source ecosystem and reproduction boundary.

As a reading path, Project 15 and this appendix should be read together. Project 15 answers "how can DataAgent be used to build an enterprise semantic BI assistant?" This appendix answers "how is that project positioned, reproduced, and governed inside the DataGallery open-source ecosystem?" Together they connect the case to the ecosystem.

## G.3 DataGallery's Role in Data Engineering

From a data engineering perspective, DataGallery's value is not to hide a complex system behind a black box. Its value is to give reproducible projects a public organizational boundary. A mature data engineering case usually needs to answer seven questions: where to obtain the project, whether the code is open source, what license applies, whether example data and configuration are reproducible, whether outputs are persisted, whether failure samples and logs can be reviewed, and how later version changes are tracked. As a public entry, DataGallery helps readers check these questions in one place.

DataGallery also reminds readers that the key assets of an open-source data-agent project are not only model-calling code. For a system such as DataAgent, the long-lived assets include semantic-layer schemas, tool configurations, database connections, execution permissions, workspace directories, runtime traces, test cases, evaluation scripts, and pre-launch gates. Only when these assets are organized can an agent move from demo capability to engineering capability.
This also means that the key assets of an open-source data-agent project are not only model-calling code. For a system such as DataAgent, the long-lived assets include semantic-layer schemas, tool configurations, database connections, execution permissions, workspace directories, runtime traces, test cases, evaluation scripts, and pre-launch gates. Only when these assets are organized can an agent move from demo capability to engineering capability.

In team collaboration, DataGallery can also serve as a cross-role communication entry. Algorithm engineers care about models, prompts, and tool choices. Data engineers care about schemas, samples, sources, quality gates, and result assets. Platform engineers care about environments, permissions, service interfaces, and runtime audit. Business users care about metric definitions, query boundaries, and result interpretation. An open repository and organization entry let these roles collaborate around the same versions, documents, and issue records instead of scattered temporary notes.

## G.4 Technical Map of the DataGallery Ecosystem

The DataGallery organization profile describes the ecosystem as a Data + AI open-source organization that reconstructs the data engineering chain through an agentic paradigm (DataGallery Contributors 2026a). In engineering terms, it can be read as a layered system rather than a single executable package. Its technical map contains four pillars.
From the public organization page and project documentation, DataGallery can be read as a Data + AI open-source practice that organizes data engineering through an agentic paradigm. In engineering terms, it can be read as a layered system rather than a single executable package. Its technical map contains four pillars.

The first pillar is **DataAgent**, which is the currently open-source execution engine. It carries NL2SQL, data analysis, feature engineering, tool calling, workspace asset persistence, and service exposure. This is the part most directly used by Project 15.

Expand All @@ -38,9 +38,9 @@ The fourth pillar is an **evaluation framework** for data-intelligence tasks. Fo

These four pillars correspond closely to the structure of this book. The semantic layer connects to the chapters on metadata, data catalogs, RAG, and data products. DataAgent connects to the chapters on tool use, multi-turn interaction, agent architecture, and project delivery. Self-evolution connects to DataOps and feedback loops. Evaluation connects to data quality, benchmark construction, acceptance gates, and reproducibility.

## G.5 DataAgent as the First Open-source Entry
## G.5 DataAgent as the Current Primary Open-source Entry

DataAgent is the open-source project that currently gives readers the most concrete way to reproduce DataGallery's engineering ideas. Its public README presents it as an enterprise Data + AI agent platform with Python 3.11+, Apache License 2.0, LangGraph integration, openJiuwen integration, GaussVector-oriented semantic retrieval support, and a versioned open-source package entry (DataGallery Contributors 2026b).
DataAgent is currently a suitable open-source project for reproducing DataGallery's engineering ideas. The public project documentation presents it as an enterprise Data + AI agent platform with Python 3.11+, Apache License 2.0, LangGraph integration, openJiuwen integration, GaussVector-oriented semantic retrieval support, and a versioned open-source package entry.

From the repository structure and documentation, DataAgent can be understood through five engineering layers.

Expand Down Expand Up @@ -94,8 +94,8 @@ Third, turn the reproduction chain into an engineering checklist. Record the ver

Fourth, continuously synchronize with repository changes. After an open-source project changes, compare configuration fields, service interfaces, dependency versions, and example paths, then update the companion reproduction notes or teaching materials.

## References
## Open-source Entries

DataGallery Contributors (2026a) DataGallery organization page. Available at: https://gitcode.com/datagallery.
- DataGallery organization entry: [https://gitcode.com/datagallery](https://gitcode.com/datagallery).

DataGallery Contributors (2026b) DataAgent source repository. Available at: https://gitcode.com/datagallery/DataAgent.
- DataAgent project repository: [https://gitcode.com/datagallery/dataagent](https://gitcode.com/datagallery/dataagent).
8 changes: 4 additions & 4 deletions docs/en/appendix_h_mindspore_note.md
Original file line number Diff line number Diff line change
Expand Up @@ -250,10 +250,10 @@ This acknowledgment explains the collaboration background and resource sources f

## References

MindFace Contributors (2026) MindFace source repository. Available at: https://github.com/mindspore-lab/mindface.
MindFace Contributors (2026) MindFace source repository. https://github.com/mindspore-lab/mindface.

MindSpore Contributors (2026a) MindSpore Documentation. Available at: https://www.mindspore.cn/view/en.
MindSpore Contributors (2026a) MindSpore Documentation. https://www.mindspore.cn/view/en.

MindSpore Contributors (2026b) MindSpore source repository. Available at: https://github.com/mindspore-ai/mindspore.
MindSpore Contributors (2026b) MindSpore source repository. https://github.com/mindspore-ai/mindspore.

MindSpore Contributors (2026c) Automatic Differentiation, MindSpore Tutorials. Available at: https://www.mindspore.cn/tutorials/en/r2.9.0/beginner/autograd.html.
MindSpore Contributors (2026c) Automatic Differentiation, MindSpore Tutorials. https://www.mindspore.cn/tutorials/en/r2.9.0/beginner/autograd.html.
Loading
Loading