docs/zh/appendix_a_tools_and_frameworks_quick_reference.md |
285 |
1 |
gebru:2021 |
Datasheets for Datasets |
Gebru T, Morgenstern J, Vecchione B, Vaughan J W, Wallach H, Daumé III H, Crawford K (2021) Datasheets for Datasets. Communications of the ACM 64(12): 86-92. |
docs/zh/appendix_a_tools_and_frameworks_quick_reference.md |
287 |
2 |
mitchell:2019 |
Model Cards for Model Reporting |
Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji I D, Gebru T (2019) Model Cards for Model Reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency, ... |
docs/zh/appendix_a_tools_and_frameworks_quick_reference.md |
289 |
3 |
pushkarna:2022 |
Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI |
Pushkarna M, Zaldivar A, Kjartansson O, Cicconi P, Chen V, Efrat A, Zou Y, Mueller J, Taly A, Ehyaei A, Karkkainen K, Marathe A, Han X, Mittal A, Schuster T, Yarmand M, Sohn H, Dwarakanath N C, McCann B (2022) Data Ca... |
docs/zh/appendix_b_compliance_and_release_checklist.md |
316 |
4 |
nist:2023 |
AI Risk Management Framework (AI RMF 1.0) |
National Institute of Standards and Technology (2023) AI Risk Management Framework (AI RMF 1.0). Available at: https://www.nist.gov/itl/ai-risk-management-framework |
docs/zh/appendix_b_compliance_and_release_checklist.md |
318 |
5 |
regulation:2024 |
|
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Available at: https://eur-lex.europa.eu/el... |
docs/zh/appendix_b_compliance_and_release_checklist.md |
320 |
6 |
mitchell:2019 |
Model Cards for Model Reporting |
Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji I D, Gebru T (2019) Model Cards for Model Reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency, ... |
docs/zh/appendix_c_cost_estimation_and_resource_templates.md |
316 |
1 |
patterson:2021 |
Carbon Emissions and Large Neural Network Training |
Patterson D, Gonzalez J, Le Q, Liang C, Munguia L, Rothchild D, So D, Texier M, Dean J (2021) Carbon Emissions and Large Neural Network Training. arXiv preprint arXiv:2104.10350. |
docs/zh/appendix_c_cost_estimation_and_resource_templates.md |
318 |
2 |
narayanan:2021 |
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM |
Narayanan D, Shoeybi M, Casper J, LeGresley P, Patwary M, Catanzaro B (2021) Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM. In: Proceedings of the International Conference for High Pe... |
docs/zh/appendix_c_cost_estimation_and_resource_templates.md |
320 |
3 |
kwon:2023 |
Efficient Memory Management for Large Language Model Serving with PagedAttention |
Kwon W, Li Z, Zhuang S, Sheng Y, Zheng L, Yu C H, Gonzalez J E, Zhang H, Stoica I (2023) Efficient Memory Management for Large Language Model Serving with PagedAttention. In: Proceedings of the ACM SIGOPS 29th Symposi... |
docs/zh/appendix_d_paper_to_implementation_guide.md |
395 |
1 |
sculley:2015 |
Hidden Technical Debt in Machine Learning Systems |
Sculley D, Holt G, Golovin D, Davydov E, Phillips T, Ebner D, Chaudhary V, Young M, Dennison D (2015) Hidden Technical Debt in Machine Learning Systems. In: Advances in Neural Information Processing Systems 28. |
docs/zh/appendix_d_paper_to_implementation_guide.md |
397 |
2 |
breck:2017 |
The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction |
Breck E, Cai S, Nielsen E, Salib M, Sculley D (2017) The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In: Proceedings of the IEEE International Conference on Big Data, pp 1123-1132. |
docs/zh/appendix_d_paper_to_implementation_guide.md |
399 |
3 |
gebru:2021 |
Datasheets for Datasets |
Gebru T, Morgenstern J, Vecchione B, Vaughan J W, Wallach H, Daum茅 III H, Crawford K (2021) Datasheets for Datasets. Communications of the ACM 64(12): 86-92. |
docs/zh/appendix_d_paper_to_implementation_guide.md |
401 |
4 |
mitchell:2019 |
Model Cards for Model Reporting |
Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji I D, Gebru T (2019) Model Cards for Model Reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency, ... |
docs/zh/appendix_d_paper_to_implementation_guide.md |
403 |
5 |
pushkarna:2022 |
Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI |
Pushkarna M, Zaldivar A, Kjartansson O, Cicconi J, Chen V, Efrat A, Zou Y, Mueller J, Taly A, Ehyaei A, Karkkainen K, Marathe A, Han X, Mittal A, Schuster T, Yarmand M, Sohn H, Dwarakanath N C, McCann B (2022) Data Ca... |
docs/zh/appendix_e_common_bug_debugging_manual.md |
447 |
1 |
sculley:2015 |
Hidden Technical Debt in Machine Learning Systems |
Sculley D, Holt G, Golovin D, Davydov E, Phillips T, Ebner D, Chaudhary V, Young M, Dennison D (2015) Hidden Technical Debt in Machine Learning Systems. In: Advances in Neural Information Processing Systems 28. |
docs/zh/appendix_e_common_bug_debugging_manual.md |
449 |
2 |
breck:2017 |
The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction |
Breck E, Cai S, Nielsen E, Salib M, Sculley D (2017) The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In: Proceedings of the IEEE International Conference on Big Data, pp 1123-1132. |
docs/zh/appendix_e_common_bug_debugging_manual.md |
451 |
3 |
amershi:2019 |
Software Engineering for Machine Learning: A Case Study |
Amershi S, Begel A, Bird C, Devanbu P, Gall H, Kamar E, Nagappan N, Nushi B, Zimmermann T (2019) Software Engineering for Machine Learning: A Case Study. In: Proceedings of the 41st International Conference on Softwar... |
docs/zh/appendix_e_common_bug_debugging_manual.md |
453 |
4 |
google:2016 |
Site Reliability Engineering: How Google Runs Production Systems |
Google SRE (2016) Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media. |
docs/zh/appendix_f_terminology_and_chinese_english_mapping.md |
401 |
1 |
mitchell:2019 |
Model Cards for Model Reporting |
Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji I D, Gebru T (2019) Model Cards for Model Reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency, ... |
docs/zh/appendix_f_terminology_and_chinese_english_mapping.md |
403 |
2 |
gebru:2021 |
Datasheets for Datasets |
Gebru T, Morgenstern J, Vecchione B, Vaughan J W, Wallach H, Daum茅 III H, Crawford K (2021) Datasheets for Datasets. Communications of the ACM 64(12): 86-92. |
docs/zh/appendix_f_terminology_and_chinese_english_mapping.md |
405 |
3 |
pushkarna:2022 |
Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI |
Pushkarna M, Zaldivar A, Kjartansson O, Cicconi J, Chen V, Efrat A, Zou Y, Mueller J, Taly A, Ehyaei A, Karkkainen K, Marathe A, Han X, Mittal A, Schuster T, Yarmand M, Sohn H, Dwarakanath N C, McCann B (2022) Data Ca... |
docs/zh/appendix_f_terminology_and_chinese_english_mapping.md |
407 |
4 |
iso:2022 |
|
ISO/IEC 22989:2022 Information technology - Artificial intelligence - Artificial intelligence concepts and terminology. |
docs/zh/appendix_g_mindspore_note.md |
51 |
1 |
mindspore:2026 |
MindSpore Documentation |
MindSpore Contributors (2026) MindSpore Documentation. Available at: https://www.mindspore.cn/view/en. |
docs/zh/appendix_g_mindspore_note.md |
53 |
2 |
mindspore:2026 |
MindSpore source repository |
MindSpore Contributors (2026) MindSpore source repository. Available at: https://github.com/mindspore-ai/mindspore. |
docs/zh/appendix_g_mindspore_note.md |
55 |
3 |
mindspore:2026 |
Automatic Differentiation, MindSpore Tutorials |
MindSpore Contributors (2026) Automatic Differentiation, MindSpore Tutorials. Available at: https://www.mindspore.cn/tutorials/en/r2.9.0/beginner/autograd.html. |
docs/zh/part1/ch01_data_change.md |
314 |
13 |
heafield:2011 |
KenLM: Faster and Smaller Language Model Queries |
Heafield K (2011) KenLM: Faster and Smaller Language Model Queries. In: Proceedings of the Sixth Workshop on Statistical Machine Translation, pp 187-197. |
docs/zh/part1/ch01_data_change.md |
316 |
14 |
broder:1997 |
On the Resemblance and Containment of Documents |
Broder A Z (1997) On the Resemblance and Containment of Documents. In: Proceedings of the Compression and Complexity of Sequences, pp 21-29. |
docs/zh/part1/ch01_data_change.md |
340 |
26 |
bloom:1970 |
Space/time Trade-offs in Hash Coding with Allowable Errors |
Bloom B H (1970) Space/time Trade-offs in Hash Coding with Allowable Errors. Communications of the ACM 13(7):422-426. |
docs/zh/part1/ch02_quality_framework.md |
488 |
1 |
cohen:1960 |
A Coefficient of Agreement for Nominal Scales |
Cohen J (1960) A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1):37-46. |
docs/zh/part1/ch02_quality_framework.md |
501 |
7 |
chen:2021 |
Evaluating Large Language Models Trained on Code (HumanEval) |
Chen M, Tworek J, Jun H, Yuan Q, Pinto H P d O, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G, others (2021) Evaluating Large Language Models Trained on Code (HumanEval). arXiv preprint arXiv:2107.03374. |
docs/zh/part1/ch02_quality_framework.md |
503 |
8 |
cobbe:2021 |
Training Verifiers to Solve Math Word Problems (GSM8K) |
Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, Plappert M, Tworek J, Hilton J, Nakano R, Hesse C, Schulman J (2021) Training Verifiers to Solve Math Word Problems (GSM8K). arXiv preprint arXiv:2110.14168. |
docs/zh/part1/ch02_quality_framework.md |
505 |
9 |
hendrycks:2021 |
Measuring Massive Multitask Language Understanding (MMLU) |
Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, Steinhardt J (2021) Measuring Massive Multitask Language Understanding (MMLU). In: International Conference on Learning Representations. |
docs/zh/part1/ch02_quality_framework.md |
507 |
10 |
broder:1997 |
On the Resemblance and Containment of Documents |
Broder A Z (1997) On the Resemblance and Containment of Documents. In: Proceedings of the Compression and Complexity of Sequences, pp 21-29. |
docs/zh/part1/ch02_quality_framework.md |
509 |
11 |
heafield:2011 |
KenLM: Faster and Smaller Language Model Queries |
Heafield K (2011) KenLM: Faster and Smaller Language Model Queries. In: Proceedings of the Sixth Workshop on Statistical Machine Translation, pp 187-197. |
docs/zh/part1/ch03_data_stack.md |
341 |
3 |
broder:1997 |
On the Resemblance and Containment of Documents |
Broder A Z (1997) On the Resemblance and Containment of Documents. In: Proceedings of the Compression and Complexity of Sequences, pp 21-29. |
docs/zh/part1/ch03_data_stack.md |
343 |
4 |
heafield:2011 |
KenLM: Faster and Smaller Language Model Queries |
Heafield K (2011) KenLM: Faster and Smaller Language Model Queries. In: Proceedings of the Sixth Workshop on Statistical Machine Translation, pp 187-197. |
docs/zh/part10/ch31_agent_architecture.md |
576 |
1 |
besta:2024 |
Graph of Thoughts: Solving Elaborate Problems with Large Language Models |
Besta M, Blach N, Kubicek A, Gerstenberger R, Podstawski M, Gianinazzi L, Gajda J, Lehmann T, Niewiadomski H, Nyczyk P, Hoefler T (2024) Graph of Thoughts: Solving Elaborate Problems with Large Language Models. In: Pr... |
docs/zh/part10/ch31_agent_architecture.md |
578 |
2 |
gao:2023 |
PAL: Program-aided Language Models |
Gao L, Madaan A, Zhou S, Alon U, Liu P, Yang Y, Callan J, Neubig G (2023) PAL: Program-aided Language Models. In: Proceedings of the 40th International Conference on Machine Learning, pp 10764-10799. |
docs/zh/part10/ch31_agent_architecture.md |
580 |
3 |
karpas:2022 |
MRKL Systems: A Modular, Neuro-Symbolic Architecture That Combines Large Language Models, External Knowledge Sources and Discrete Reasoning |
Karpas E, Abend O, Belinkov Y, Lenz B, Lieber O, Ratner N, Shoham Y, Bata H, Levine Y, Leyton-Brown K, Muhlgay D, Rozen N, Schwartz E, Shashua A, Shuster K, Tenenbaum J, Wolf L, Zettlemoyer L, Riedel S (2022) MRKL Sys... |
docs/zh/part10/ch31_agent_architecture.md |
582 |
4 |
kreuzberger:2023 |
Machine Learning Operations (MLOps): Overview, Definition, and Architecture |
Kreuzberger D, Kühl N, Hirschl S (2023) Machine Learning Operations (MLOps): Overview, Definition, and Architecture. IEEE Access 11:31866-31879. |
docs/zh/part10/ch31_agent_architecture.md |
584 |
5 |
madaan:2023 |
Self-Refine: Iterative Refinement with Self-Feedback |
Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, Alon U, Dziri N, Prabhumoye S, Yang Y, Gupta S, Majumder B P, Hermann K, Welleck S, Yazdanbakhsh A, Clark P (2023) Self-Refine: Iterative Refinement with Se... |
docs/zh/part10/ch31_agent_architecture.md |
586 |
6 |
mialon:2023 |
Augmented Language Models: A Survey |
Mialon G, Dessì R, Lomeli M, Nalmpantis C, Pasunuru R, Raileanu R, Rozière B, Schick T, Dwivedi-Yu J, Celikyilmaz A, Grave E, LeCun Y, Scialom T (2023) Augmented Language Models: A Survey. Transactions on Machine Lear... |
docs/zh/part10/ch31_agent_architecture.md |
588 |
7 |
nakano:2021 |
WebGPT: Browser-assisted question-answering with human feedback |
Nakano R, Hilton J, Balaji S, Wu J, Ouyang L, Kim C, Hesse C, Jain S, Kosaraju V, Saunders W, Jiang X, Cobbe K, Eloundou T, Krueger G, Button K, Knight M, Chess B, Schulman J (2021) WebGPT: Browser-assisted question-a... |
docs/zh/part10/ch31_agent_architecture.md |
590 |
8 |
nist:2023 |
Artificial Intelligence Risk Management Framework (AI RMF 1.0) |
NIST (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. |
docs/zh/part10/ch31_agent_architecture.md |
592 |
9 |
nist:2024 |
Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile |
NIST (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. |
docs/zh/part10/ch31_agent_architecture.md |
594 |
10 |
park:2023 |
Generative Agents: Interactive Simulacra of Human Behavior |
Park J S, O'Brien J C, Cai C J, Morris M R, Liang P, Bernstein M S (2023) Generative Agents: Interactive Simulacra of Human Behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Tec... |
docs/zh/part10/ch31_agent_architecture.md |
596 |
11 |
patil:2023 |
Gorilla: Large Language Model Connected with Massive APIs |
Patil S G, Zhang T, Wang X, Gonzalez J E (2023) Gorilla: Large Language Model Connected with Massive APIs. arXiv preprint arXiv:2305.15334. |
docs/zh/part10/ch31_agent_architecture.md |
598 |
12 |
qin:2024 |
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs |
Qin Y, Liang S, Ye Y, Zhu K, Yan L, Lu Y, Lin Y, Cong X, Tang X, Qian B, Zhao S, Tian R, Xie R, Zhou J, Gerstein M, Li D, Liu Z, Sun M (2024) ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world API... |
docs/zh/part10/ch31_agent_architecture.md |
600 |
13 |
schick:2023 |
Toolformer: Language Models Can Teach Themselves to Use Tools |
Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, Lomeli M, Hambro E, Zettlemoyer L, Cancedda N, Scialom T (2023) Toolformer: Language Models Can Teach Themselves to Use Tools. In: Advances in Neural Information Processing... |
docs/zh/part10/ch31_agent_architecture.md |
602 |
14 |
shinn:2023 |
Reflexion: Language Agents with Verbal Reinforcement Learning |
Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S (2023) Reflexion: Language Agents with Verbal Reinforcement Learning. In: Advances in Neural Information Processing Systems 36. |
docs/zh/part10/ch31_agent_architecture.md |
604 |
15 |
wang:2023 |
A Survey on Large Language Model based Autonomous Agents |
Wang L, Ma C, Feng X, Zhang Z, Yang H, Zhang J, Chen Z, Tang J, Chen X, Lin Y, Zhao W X, Wei Z, Wen J-R (2023) A Survey on Large Language Model based Autonomous Agents. arXiv preprint arXiv:2308.11432. |
docs/zh/part10/ch31_agent_architecture.md |
606 |
16 |
wu:2023 |
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation |
Wu Q, Bansal G, Zhang J, Wu Y, Li B, Zhu E, Jiang L, Zhang X, Zhang S, Liu J, Awadallah A H, White R W, Burger D, Wang C (2023) AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv preprint ... |
docs/zh/part10/ch31_agent_architecture.md |
608 |
17 |
xi:2023 |
The Rise and Potential of Large Language Model Based Agents: A Survey |
Xi Z, Chen W, Guo X, He W, Ding Y, Hong B, Zhang M, Wang J, Jin S, Zhou E, Zheng R, Fan X, Wang X, Xiong L, Zhou Y, Wang W, Jiang C, Zou Y, Liu X, Yin Z, Dou S, Weng R, Cheng W, Zhang Q, Qin W, Zheng Y, Qiu X, Huang X... |
docs/zh/part10/ch31_agent_architecture.md |
610 |
18 |
yao:2023 |
ReAct: Synergizing Reasoning and Acting in Language Models |
Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, Cao Y (2023) ReAct: Synergizing Reasoning and Acting in Language Models. In: International Conference on Learning Representations. |
docs/zh/part10/ch31_agent_architecture.md |
612 |
19 |
yao:2023 |
Tree of Thoughts: Deliberate Problem Solving with Large Language Models |
Yao S, Yu D, Zhao J, Shafran I, Griffiths T L, Cao Y, Narasimhan K (2023) Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In: Advances in Neural Information Processing Systems 36. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
415 |
1 |
barbaresi:2021 |
Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction |
Barbaresi A (2021) Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, pp 122-131. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
417 |
2 |
blecher:2023 |
Nougat: Neural Optical Understanding for Academic Documents |
Blecher N, Cresci G, Ballas N, Bautista M (2023) Nougat: Neural Optical Understanding for Academic Documents. arXiv preprint arXiv:2308.13418. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
419 |
3 |
carlini:2021 |
Extracting Training Data from Large Language Models |
Carlini N, Tramer F, Wallace E, Jagielski M, Herbert-Voss A, Lee K, Roberts A, Brown T, Song D, Erlingsson U, Oprea A, Raffel C (2021) Extracting Training Data from Large Language Models. In: Proceedings of the 30th U... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
421 |
4 |
chen:2024 |
Data-Juicer: A One-Stop Data Processing System for Large Language Models |
Chen J, Yan X, Lin D, Qu X, Wang Y, Huang X, Zhao Z, Yu T, Zhang Z, Li H, Zheng Y, Xu R, Zhu J, Qiu X (2024) Data-Juicer: A One-Stop Data Processing System for Large Language Models. In: Proceedings of the ACM SIGMOD ... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
423 |
5 |
chowdhery:2022 |
PaLM: Scaling Language Modeling with Pathways |
Chowdhery A, Narang S, Devlin J, Bosma M, Mishra G, Roberts A, Barham P, Chung H W, Sutton C, Gehrmann S, Schuh P, Shi K, Tsvyashchenko S, Maynez J, Rao A, Barnes P, Tay Y, Shazeer N, Prabhakaran V, Reif E, Du N, Hutc... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
425 |
6 |
dodge:2021 |
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus |
Dodge J, Sap M, Marasović A, Agnew W, Ilharco G, Groeneveld D, Mitchell M, Gardner M (2021) Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In: Proceedings of the 2021 Conference ... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
427 |
7 |
gao:2020 |
The Pile: An 800GB Dataset of Diverse Text for Language Modeling |
Gao L, Biderman S, Black S, Golding L, Hoppe T, Foster C, Phang J, He H, Thite A, Nabeshima N, Presser S, Leahy C (2020) The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv preprint arXiv:2101.00027. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
429 |
8 |
huang:2022 |
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking |
Huang Y, Lv T, Cui L, Lu Y, Wei F (2022) LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking. In: Proceedings of the 30th ACM International Conference on Multimedia, pp 4083-4091. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
431 |
9 |
kim:2022 |
OCR-free Document Understanding Transformer |
Kim G, Hong T, Yim M, Nam J, Park J, Yim J, Hwang W, Yun S, Han D, Park S (2022) OCR-free Document Understanding Transformer. In: European Conference on Computer Vision, pp 498-517. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
433 |
10 |
laurencon:2022 |
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset |
Laurençon H, Saulnier L, Wang T, Akiki C, del Moral A V, Le Scao T, Von Werra L, Mou C, González Ponferrada E, Nguyen H, Frohberg J, Šaško M, Lhoest Q, McMillan-Major A, Dupont G, Biderman S, Rogers A, Allal L B, De T... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
435 |
11 |
lee:2022 |
Deduplicating Training Data Makes Language Models Better |
Lee K, Ippolito D, Nystrom A, Zhang C, Eck D, Callison-Burch C, Carlini N (2022) Deduplicating Training Data Makes Language Models Better. In: Proceedings of the 60th Annual Meeting of the Association for Computationa... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
437 |
12 |
longpre:2023 |
The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing and Attribution in AI |
Longpre S, Mahari R, Lee A, et al. (2023) The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing and Attribution in AI. arXiv preprint arXiv:2310.16787. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
439 |
13 |
nguyen:2024 |
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages |
Nguyen T, et al. (2024) CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lang... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
441 |
14 |
ortiz:2020 |
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages |
Ortiz Suárez P J, Sagot B, Romary L (2020) A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages. In: Proceedings of the 12th Language Resources and Evaluation Conference, pp 1703-1714. |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
443 |
15 |
pfitzmann:2022 |
DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis |
Pfitzmann B, Auer C, Dolfi M, Nassar A S, Staar P (2022) DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis. In: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Minin... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
445 |
16 |
penedo:2024 |
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale |
Penedo G, Kydlíček H, Allal L B, Lozhkov A, Mitchell M, Raffel C, von Werra L, Wolf T (2024) The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. In: Advances in Neural Information Processing Sys... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
447 |
17 |
penedo:2023 |
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only |
Penedo G, Malartic Q, Hesslow D, Cojocaru R, Cappelli A, Alobeidli H, Pannier B, Almazrouei E, Launay J (2023) The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only. In: Advances in N... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
449 |
18 |
raffel:2020 |
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer |
Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu P J (2020) Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21(140):1... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
451 |
19 |
soldaini:2024 |
Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research |
Soldaini L, Kinney R, Bhagia A, Schwenk D, Atkinson D, Authur A, Bogin B, Chen X, Dumas G, Elazar Y, Hofmann V, Jha A H, Kumar S, Lucy L, Lyu X, Lambert N, Magnusson I, Morrison J, Muennighoff N, Naik A, Nam G, Peters... |
docs/zh/part10/ch32_auto_collection_parsing_cleaning.md |
453 |
20 |
wenzek:2020 |
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data |
Wenzek G, Lachaux M-A, Conneau A, Chaudhary V, Guzmán F, Joulin A, Grave E (2020) CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. In: Proceedings of the 12th Language Resources and Evaluation ... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
432 |
1 |
alemohammad:2024 |
Self-Consuming Generative Models Go MAD |
Alemohammad S, Casco-Rodriguez J, Luzi L, et al. (2024) Self-Consuming Generative Models Go MAD. In: International Conference on Learning Representations. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
434 |
2 |
bai:2022 |
Constitutional AI: Harmlessness from AI Feedback |
Bai Y, Kadavath S, Kundu S, et al. (2022) Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
436 |
3 |
cui:2023 |
UltraFeedback: Boosting Language Models with Scaled AI Feedback |
Cui G, Yuan L, Ding N, Yao G, Zhu W, Ni Y, Xie G, Liu Z, Sun M (2023) UltraFeedback: Boosting Language Models with Scaled AI Feedback. arXiv preprint arXiv:2310.01377. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
438 |
4 |
dubois:2023 |
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback |
Dubois Y, Li X, Taori R, Zhang T, Gulrajani I, Ba J, Guestrin C, Liang P, Hashimoto T B (2023) AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. In: Advances in Neural Information Processi... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
440 |
5 |
gerstgrasser:2024 |
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data |
Gerstgrasser M, Schaeffer R, Dey A, et al. (2024) Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv preprint arXiv:2404.01413. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
442 |
6 |
kim:2024 |
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models |
Kim S, Shin J, Cho Y, Jang J, Longpre S, Lee H, Yun S, Shin S, Kim S, Thorne J, Seo M (2024) Prometheus: Inducing Fine-grained Evaluation Capability in Language Models. In: International Conference on Learning Represe... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
444 |
7 |
kim:2024 |
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models |
Kim S, Suk J, Longpre S, Lin B Y, Shin J, Welleck S, Neubig G, Lee M, Lee K, Seo M (2024) Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. arXiv preprint arXiv:2405.01535. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
446 |
8 |
koh:2021 |
WILDS: A Benchmark of in-the-Wild Distribution Shifts |
Koh P W, Sagawa S, Marklund H, et al. (2021) WILDS: A Benchmark of in-the-Wild Distribution Shifts. In: Proceedings of the 38th International Conference on Machine Learning, pp 5637-5664. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
448 |
9 |
lambert:2024 |
RewardBench: Evaluating Reward Models for Language Modeling |
Lambert N, Pyatkin V, Morrison J, Miranda L, Lin B Y, Chandu K, Dziri N, Kumar S, Zick T, Choi Y, Smith N A, Hajishirzi H (2024) RewardBench: Evaluating Reward Models for Language Modeling. arXiv preprint arXiv:2403.1... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
450 |
10 |
liang:2023 |
Holistic Evaluation of Language Models |
Liang P, Bommasani R, Lee T, et al. (2023) Holistic Evaluation of Language Models. Transactions on Machine Learning Research. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
452 |
11 |
liu:2023 |
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment |
Liu Y, Iter D, Xu Y, et al. (2023) G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp 2511-2522. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
454 |
12 |
lin:2024 |
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild |
Lin B Y, et al. (2024) WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild. arXiv preprint arXiv:2406.04770. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
456 |
13 |
ouyang:2022 |
Training language models to follow instructions with human feedback |
Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R (2022) Training ... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
458 |
14 |
perez:2022 |
Red Teaming Language Models with Language Models |
Perez E, Huang S, Song F, Cai T, Ring R, Aslanides J, Glaese A, McAleese N, Irving G (2022) Red Teaming Language Models with Language Models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Lang... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
460 |
15 |
rafailov:2023 |
Direct Preference Optimization: Your Language Model is Secretly a Reward Model |
Rafailov R, Sharma A, Mitchell E, Manning C D, Ermon S, Finn C (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In: Advances in Neural Information Processing Systems 36. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
462 |
16 |
ribeiro:2020 |
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList |
Ribeiro M T, Wu T, Guestrin C, Singh S (2020) Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp 4902-4912. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
464 |
17 |
shumailov:2024 |
AI models collapse when trained on recursively generated data |
Shumailov I, Shumaylov Z, Zhao Y, et al. (2024) AI models collapse when trained on recursively generated data. Nature 631:755-759. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
466 |
18 |
wang:2023 |
Self-Instruct: Aligning Language Models with Self-Generated Instructions |
Wang Y, Kordi Y, Mishra S, et al. (2023) Self-Instruct: Aligning Language Models with Self-Generated Instructions. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pp 13484-... |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
468 |
19 |
zheng:2023 |
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena |
Zheng L, Chiang W-L, Sheng Y, et al. (2023) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In: Advances in Neural Information Processing Systems 36. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
470 |
20 |
zhu:2023 |
JudgeLM: Fine-tuned Large Language Models are Scalable Judges |
Zhu L, Wang X, Wang Y, et al. (2023) JudgeLM: Fine-tuned Large Language Models are Scalable Judges. arXiv preprint arXiv:2310.17631. |
docs/zh/part10/ch33_labeling_synthesis_evaluation.md |
472 |
21 |
zhou:2023 |
LIMA: Less Is More for Alignment |
Zhou C, Liu P, Xu P, et al. (2023) LIMA: Less Is More for Alignment. In: Advances in Neural Information Processing Systems 36. |
docs/zh/part10/ch34_dataops_agent.md |
468 |
1 |
amershi:2019 |
Software Engineering for Machine Learning: A Case Study |
Amershi S, Begel A, Bird C, Devanbu P, Gall H, Kamar E, Nagappan N, Nushi B, Zimmermann T (2019) Software Engineering for Machine Learning: A Case Study. In: Proceedings of the 41st International Conference on Softwar... |
docs/zh/part10/ch34_dataops_agent.md |
470 |
2 |
breck:2019 |
Data Validation for Machine Learning |
Breck E, Polyzotis N, Roy S, Whang S E, Zinkevich M (2019) Data Validation for Machine Learning. In: Proceedings of Machine Learning and Systems 1, pp 334-347. |
docs/zh/part10/ch34_dataops_agent.md |
472 |
3 |
dang:2019 |
AIOps: Real-World Challenges and Research Innovations |
Dang Y, Lin Q, Huang P (2019) AIOps: Real-World Challenges and Research Innovations. In: Proceedings of the 41st International Conference on Software Engineering: Companion Proceedings, pp 4-5. |
docs/zh/part10/ch34_dataops_agent.md |
474 |
4 |
he:2021 |
A Survey on Automated Log Analysis for Reliability Engineering |
He S, He P, Chen Z, Yang T, Su Y, Lyu M R (2021) A Survey on Automated Log Analysis for Reliability Engineering. ACM Computing Surveys 54(6):1-37. |
docs/zh/part10/ch34_dataops_agent.md |
476 |
5 |
huyen:2022 |
Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications |
Huyen C (2022) Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications. O'Reilly Media. |
docs/zh/part10/ch34_dataops_agent.md |
478 |
6 |
kreuzberger:2023 |
Machine Learning Operations (MLOps): Overview, Definition, and Architecture |
Kreuzberger D, Kühl N, Hirschl S (2023) Machine Learning Operations (MLOps): Overview, Definition, and Architecture. IEEE Access 11:31866-31879. |
docs/zh/part10/ch34_dataops_agent.md |
480 |
7 |
makinen:2021 |
Who Needs MLOps: What Data Scientists Seek to Accomplish and How Can MLOps Help? In: Proceedings of the 2021 IEEE/ACM 1st Workshop on AI Engineering - Software Engineering for AI, pp 109-112 |
Makinen S, Skogstrom H, Laaksonen E, Mikkonen T (2021) Who Needs MLOps: What Data Scientists Seek to Accomplish and How Can MLOps Help? In: Proceedings of the 2021 IEEE/ACM 1st Workshop on AI Engineering - Software En... |
docs/zh/part10/ch34_dataops_agent.md |
482 |
8 |
lwakatare:2020 |
Large-scale Machine Learning Systems in Real-world Industrial Settings: A Review of Challenges and Solutions |
Lwakatare L E, Raj A, Crnkovic I, Bosch J, Olsson H H (2020) Large-scale Machine Learning Systems in Real-world Industrial Settings: A Review of Challenges and Solutions. Information and Software Technology 127:106368. |
docs/zh/part10/ch34_dataops_agent.md |
484 |
9 |
nist:2023 |
Artificial Intelligence Risk Management Framework (AI RMF 1.0) |
NIST (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. |
docs/zh/part10/ch34_dataops_agent.md |
486 |
10 |
nist:2024 |
Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile |
NIST (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. |
docs/zh/part10/ch34_dataops_agent.md |
488 |
11 |
paleyes:2022 |
Challenges in Deploying Machine Learning: A Survey of Case Studies |
Paleyes A, Urma R-G, Lawrence N D (2022) Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys 55(6):1-29. |
docs/zh/part10/ch34_dataops_agent.md |
490 |
12 |
sambasivan:2021 |
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI |
Sambasivan N, Kapania S, Highfill H, Akrong D, Paritosh P, Aroyo L M (2021) "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI. In: Proceedings of the 2021 CHI Conference on Huma... |
docs/zh/part10/ch34_dataops_agent.md |
492 |
13 |
tamburri:2020 |
Sustainable MLOps: Trends and Challenges |
Tamburri D A (2020) Sustainable MLOps: Trends and Challenges. In: Proceedings of the 22nd International Symposium on Symbolic and Numeric Algorithms for Scientific Computing, pp 17-23. |
docs/zh/part10/ch34_dataops_agent.md |
494 |
14 |
testi:2022 |
MLOps: A Taxonomy and a Methodology |
Testi M, Ballabio M, Frontoni E, Iannello G, Moccia S, Soda P, Vessio G (2022) MLOps: A Taxonomy and a Methodology. IEEE Access 10:63606-63618. |
docs/zh/part10/ch34_dataops_agent.md |
496 |
15 |
treveil:2020 |
Introducing MLOps: How to Scale Machine Learning in the Enterprise |
Treveil M, Omont N, Stenac C, Lefevre K, Phan D, Zentici J, Lavoillotte A, Miyazaki M, Heidmann L (2020) Introducing MLOps: How to Scale Machine Learning in the Enterprise. O'Reilly Media. |
docs/zh/part10/ch34_dataops_agent.md |
498 |
16 |
vela:2022 |
Temporal quality degradation in AI models |
Vela D, Sharp A, Zhang R, Nguyen T, Hoang A, Pianykh O S (2022) Temporal quality degradation in AI models. Scientific Reports 12:11654. |
docs/zh/part10/ch34_dataops_agent.md |
500 |
17 |
zhu:2019 |
Tools and Benchmarks for Automated Log Parsing |
Zhu J, He S, Liu J, He P, Xie Q, Zheng Z, Lyu M R (2019) Tools and Benchmarks for Automated Log Parsing. In: Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice, ... |
docs/zh/part10/ch35_security_permission_collaboration.md |
466 |
1 |
andriushchenko:2024 |
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks |
Andriushchenko M, Croce F, Flammarion N, Hein M (2024) Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks. arXiv preprint arXiv:2404.02151. |
docs/zh/part10/ch35_security_permission_collaboration.md |
468 |
2 |
chen:2024 |
StruQ: Defending Against Prompt Injection with Structured Queries |
Chen S, Piet J, Sitawarin C, Wagner D (2024) StruQ: Defending Against Prompt Injection with Structured Queries. arXiv preprint arXiv:2402.06363. |
docs/zh/part10/ch35_security_permission_collaboration.md |
470 |
3 |
debenedetti:2024 |
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents |
Debenedetti E, Zhang J, Balunović M, et al. (2024) AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In: Advances in Neural Information Processing Systems 37. |
docs/zh/part10/ch35_security_permission_collaboration.md |
472 |
4 |
ganguli:2022 |
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned |
Ganguli D, Lovitt L, Kernion J, Askell A, Bai Y, Kadavath S, Mann B, Perez E, Schiefer N, Ndousse K, Jones A, Bowman S R, Chen A, Conerly T, DasSarma N, Drain D, Elhage N, El-Showk S, Fort S, Hatfield-Dodds Z, Henigha... |
docs/zh/part10/ch35_security_permission_collaboration.md |
474 |
5 |
greshake:2023 |
Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection |
Greshake K, Abdelnabi S, Mishra S, et al. (2023) Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In: Proceedings of the 16th ACM Workshop on Artificia... |
docs/zh/part10/ch35_security_permission_collaboration.md |
476 |
6 |
hendrycks:2021 |
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization |
Hendrycks D, Mazeika M, Zou A, Patel S, Zhu C, Navarro J, Mu J, Song D, Li B, Steinhardt J (2021) The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization. In: Proceedings of the IEEE/CV... |
docs/zh/part10/ch35_security_permission_collaboration.md |
478 |
7 |
huang:2024 |
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation |
Huang Y, Gupta S, Xia M, Li K, Chen D (2024) Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation. In: International Conference on Learning Representations. |
docs/zh/part10/ch35_security_permission_collaboration.md |
480 |
8 |
lapid:2023 |
Open Sesame! Universal Black Box Jailbreaking of Large Language Models |
Lapid R, Langberg R, Sipper M (2023) Open Sesame! Universal Black Box Jailbreaking of Large Language Models. arXiv preprint arXiv:2309.01446. |
docs/zh/part10/ch35_security_permission_collaboration.md |
482 |
9 |
liu:2023 |
Prompt Injection Attack against LLM-Integrated Applications |
Liu Y, Deng G, Li Y, et al. (2023) Prompt Injection Attack against LLM-Integrated Applications. arXiv preprint arXiv:2306.05499. |
docs/zh/part10/ch35_security_permission_collaboration.md |
484 |
10 |
perez:2022 |
Red Teaming Language Models with Language Models |
Perez E, Huang S, Song F, Cai T, Ring R, Aslanides J, Glaese A, McAleese N, Irving G (2022) Red Teaming Language Models with Language Models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Lang... |
docs/zh/part10/ch35_security_permission_collaboration.md |
486 |
11 |
ruan:2024 |
Identifying the Risks of LM Agents with an LM-Emulated Sandbox |
Ruan Y, Dong H, Wang A, Pitis S, Zhou Y, Ba J, Dubois Y, Maddison C J, Hashimoto T B (2024) Identifying the Risks of LM Agents with an LM-Emulated Sandbox. In: International Conference on Learning Representations. |
docs/zh/part10/ch35_security_permission_collaboration.md |
488 |
12 |
tian:2023 |
Evil Geniuses: Delving into the Safety of LLM-based Agents |
Tian Y, Yang X, Zhang J, Dong Y, Su H (2023) Evil Geniuses: Delving into the Safety of LLM-based Agents. arXiv preprint arXiv:2311.11855. |
docs/zh/part10/ch35_security_permission_collaboration.md |
490 |
13 |
toyer:2024 |
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game |
Toyer S, Watkins O, Mendes E A, Svegliato J, Bailey L, Wang T, Ong I, Elmaaroufi K, Abbeel P, Darrell T, Ritter A, Russell S (2024) Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game. In: Interna... |
docs/zh/part10/ch35_security_permission_collaboration.md |
492 |
14 |
wallace:2024 |
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions |
Wallace E, Xiao K, Leike R, Weng L, Heidecke J, Beutel A (2024) The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. arXiv preprint arXiv:2404.13208. |
docs/zh/part10/ch35_security_permission_collaboration.md |
494 |
15 |
wei:2023 |
Jailbroken: How Does LLM Safety Training Fail? |
Wei A, Haghtalab N, Steinhardt J (2023) Jailbroken: How Does LLM Safety Training Fail? arXiv preprint arXiv:2307.02483. |
docs/zh/part10/ch35_security_permission_collaboration.md |
496 |
16 |
yi:2023 |
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models |
Yi J, Xie Y, Zhu B, Hines K, Kiciman E, Sun G, Xie X, Wu F (2023) Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models. arXiv preprint arXiv:2312.14197. |
docs/zh/part10/ch35_security_permission_collaboration.md |
498 |
17 |
zhan:2024 |
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents |
Zhan Q, Liang Z, Ying Z, Kang D (2024) InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. In: Findings of the Association for Computational Linguistics: ACL 2024, pp 10... |
docs/zh/part10/ch35_security_permission_collaboration.md |
500 |
18 |
zou:2023 |
Universal and Transferable Adversarial Attacks on Aligned Language Models |
Zou A, Wang Z, Carlini N, Nasr M, Kolter J Z, Fredrikson M (2023) Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv preprint arXiv:2307.15043. |
docs/zh/part11/ch36_compliance_framework_and_governance.md |
1180 |
5 |
european:2022 |
Data Protection Engineering |
European Union Agency for Cybersecurity (ENISA) (2022) Data Protection Engineering. ENISA Report. |
docs/zh/part11/ch36_compliance_framework_and_governance.md |
1186 |
8 |
hoepman:2014 |
Privacy Design Strategies |
Hoepman J-H (2014) Privacy Design Strategies. In IFIP International Information Security Conference, pp 446-459. |
docs/zh/part11/ch36_compliance_framework_and_governance.md |
1190 |
10 |
dwork:2008 |
Differential Privacy: A Survey of Results |
Dwork C (2008) Differential Privacy: A Survey of Results. In Theory and Applications of Models of Computation, Springer Berlin Heidelberg, pp 1-19. |
docs/zh/part11/ch37_federated_learning_and_privacy_preserving_technologies.md |
448 |
4 |
dwork:2011 |
Differential Privacy |
Dwork C (2011) Differential Privacy. In Encyclopedia of Cryptography and Security, Springer US, pp 338-340. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
354 |
1 |
bai:2025 |
Qwen2.5-VL Technical Report |
Bai, S., Chen, K., Liu, X., et al. (2025). Qwen2.5-VL Technical Report. arXiv preprint arXiv:2502.13923. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
356 |
2 |
blecher:2023 |
Nougat: Neural Optical Understanding for Academic Documents |
Blecher, L., Cucurull, G., Scialom, T., and Stojnic, R. (2023). Nougat: Neural Optical Understanding for Academic Documents. arXiv preprint arXiv:2308.13418. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
358 |
3 |
huang:2022 |
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking. Proc |
Huang, Y., Lv, T., Cui, L., Lu, Y., and Wei, F. (2022). LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking. Proc. ACM Multimedia. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
360 |
4 |
huang:2019 |
ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction |
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., and Jawahar, C.V. (2019). ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction. Proc. ICDAR, pp. 1516–1520. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
362 |
5 |
hu:2021 |
LoRA: Low-Rank Adaptation of Large Language Models |
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv preprint arXiv:2106.09685. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
364 |
6 |
jaume:2019 |
FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents |
Jaume, G., Ekenel, H.K., and Thiran, J.-P. (2019). FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents. ICDAR Workshop. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
366 |
7 |
kuhn:1955 |
The Hungarian Method for the Assignment Problem |
Kuhn, H.W. (1955). The Hungarian Method for the Assignment Problem. Naval Research Logistics Quarterly, 2(1–2), pp. 83–97. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
368 |
8 |
levenshtein:1965 |
Binary Codes Capable of Correcting Deletions, Insertions and Reversals |
Levenshtein, V.I. (1965). Binary Codes Capable of Correcting Deletions, Insertions and Reversals. Soviet Physics Doklady, 10, pp. 707–710. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
370 |
9 |
liu:2024 |
A Survey on Hallucination in Large Vision-Language Models |
Liu, H., Xue, W., Chen, Y., et al. (2024). A Survey on Hallucination in Large Vision-Language Models. arXiv preprint arXiv:2402.00253. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
372 |
10 |
mathew:2021 |
DocVQA: A Dataset for VQA on Document Images |
Mathew, M., Karatzas, D., and Jawahar, C.V. (2021). DocVQA: A Dataset for VQA on Document Images. Proc. WACV. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
374 |
11 |
niu:2025 |
MinerU 2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing |
Niu, J., Liu, Z., Gu, Z., et al. (2025). MinerU 2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing. arXiv preprint. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
376 |
12 |
park:2019 |
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing |
Park, S., Shin, S., Lee, B., et al. (2019). CORD: A Consolidated Receipt Dataset for Post-OCR Parsing. NeurIPS Workshop on Document Intelligence. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
378 |
13 |
rafailov:2024 |
Direct Preference Optimization: Your Language Model Is Secretly a Reward Model |
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D., and Finn, C. (2024). Direct Preference Optimization: Your Language Model Is Secretly a Reward Model. Proc. NeurIPS. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
380 |
14 |
schulman:2017 |
Proximal Policy Optimization Algorithms |
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
382 |
15 |
shao:2024 |
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models |
Shao, Z., Wang, P., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
384 |
16 |
tianchi:2022 |
CHIP 2022 Shared Task: Medical Invoice OCR Element Extraction Dataset |
Tianchi, A. and CHIP Committee (2022). CHIP 2022 Shared Task: Medical Invoice OCR Element Extraction Dataset. Aliyun Tianchi Platform. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
386 |
17 |
xu:2020 |
LayoutLM: Pre-training of Text and Layout for Document Image Understanding. Proc |
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., and Zhou, M. (2020). LayoutLM: Pre-training of Text and Layout for Document Image Understanding. Proc. ACM SIGKDD, pp. 1192–1200. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
388 |
18 |
xue:2021 |
TGRNet: A Table Graph Reconstruction Network for Table Structure Recognition |
Xue, W., Yu, B., Wang, W., Tao, D., and Li, Q. (2021). TGRNet: A Table Graph Reconstruction Network for Table Structure Recognition. arXiv preprint arXiv:2106.10598. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
390 |
19 |
yang:2023 |
Modeling Entities as Semantic Points for Visual Information Extraction in the Wild |
Yang, Z., Long, R., Wang, P., et al. (2023). Modeling Entities as Semantic Points for Visual Information Extraction in the Wild. Proc. CVPR. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
392 |
20 |
zhang:2022 |
CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark |
Zhang, N., Chen, M., Bi, Z., et al. (2022). CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark. Proc. ACL, pp. 7888–7915. |
docs/zh/part12/ch38_structbill_cn_dataset.md |
394 |
21 |
zhong:2020 |
Image-based Table Recognition: Data, Model, and Evaluation |
Zhong, X., ShafieiBavani, E., and Jimeno Yepes, A. (2020). Image-based Table Recognition: Data, Model, and Evaluation. arXiv preprint arXiv:2011.13534. |
docs/zh/part12/ch39_sparse_table_bench_dataset.md |
280 |
1 |
zhong:2020 |
Image-based Table Recognition: Data, Model, and Evaluation |
1. Zhong, X., ShafieiBavani, E., & Yepes, A. J. (2020). Image-based Table Recognition: Data, Model, and Evaluation. ECCV 2020. |
docs/zh/part12/ch39_sparse_table_bench_dataset.md |
281 |
2 |
smock:2022 |
PubTables-1M: Towards Comprehensive Table Extraction From Unstructured Documents |
2. Smock, B., Pesala, R., & Abraham, R. (2022). PubTables-1M: Towards Comprehensive Table Extraction From Unstructured Documents. CVPR 2022. |
docs/zh/part12/ch39_sparse_table_bench_dataset.md |
282 |
3 |
zhu:2021 |
TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance |
3. Zhu, F., Lei, W., Huang, Y., Wang, C., Zhang, S., Lv, J., Feng, F., & Chua, T.-S. (2021). TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance. ACL 2021. |
docs/zh/part12/ch39_sparse_table_bench_dataset.md |
283 |
4 |
pandas:2026 |
pandas Documentation |
4. Pandas Development Team. (2026). pandas Documentation. https://pandas.pydata.org/docs/ |
docs/zh/part12/ch39_sparse_table_bench_dataset.md |
284 |
5 |
apache:2026 |
Apache Arrow Documentation |
5. Apache Arrow Contributors. (2026). Apache Arrow Documentation. https://arrow.apache.org/docs/ |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
217 |
1 |
masry:2022 |
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning |
1. Masry, A., Long, D. X., Tan, J. Q., Joty, S., & Hoque, E. (2022). ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. ACL 2022. |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
218 |
2 |
methani:2020 |
PlotQA: Reasoning over Scientific Plots |
2. Methani, N., Ganguly, P., Khapra, M. M., & Kumar, P. (2020). PlotQA: Reasoning over Scientific Plots. WACV 2020. |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
219 |
3 |
kahou:2017 |
FigureQA: An Annotated Figure Dataset for Visual Reasoning |
3. Kahou, S. E., Michalski, V., Atkinson, A., Kádár, Á., Trischler, A., & Bengio, Y. (2017). FigureQA: An Annotated Figure Dataset for Visual Reasoning. arXiv:1710.07300. |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
220 |
4 |
kafle:2018 |
DVQA: Understanding Data Visualizations via Question Answering |
4. Kafle, K., Price, B., Cohen, S., & Kanan, C. (2018). DVQA: Understanding Data Visualizations via Question Answering. CVPR 2018. |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
221 |
5 |
mathew:2021 |
DocVQA: A Dataset for VQA on Document Images |
5. Mathew, M., Karatzas, D., & Jawahar, C. V. (2021). DocVQA: A Dataset for VQA on Document Images. WACV 2021. |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
222 |
6 |
masry:2025 |
|
6. Masry, A., Islam, M. S., Ahmed, M., Bajaj, A., Kabir, F., Kartha, A., ... & Joty, S. (2025, July). Chartqapro: A more diverse and challenging benchmark for chart question answering. In Findings of the Association f... |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
223 |
7 |
xie:2026 |
Infochartqa: A benchmark for multimodal question answering on infographic charts |
7. Xie, T., Lin, M., Liu, M., Ye, Y., Chen, C., & Liu, S. (2026). Infochartqa: A benchmark for multimodal question answering on infographic charts. Advances in Neural Information Processing Systems, 38. |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
224 |
8 |
foroutan:2025 |
|
8. Foroutan, N., Romanou, A., Ansaripour, M., Eisenschlos, J. M., Aberer, K., & Lebret, R. (2025, July). Wikimixqa: a multimodal benchmark for question answering over tables and charts. In Findings of the Association ... |
docs/zh/part12/ch40_multi_chart_infographic_reasoning_dataset.md |
225 |
9 |
zhu:2025 |
|
9. Zhu, Z., Jia, M., Zhang, Z., Li, L., & Jiang, M. (2025, April). MultiChartQA: Benchmarking vision-language models on multi-chart problems. In Proceedings of the 2025 Conference of the Nations of the Americas Chapte... |
docs/zh/part12/ch42_voice_style_control_dataset.md |
390 |
1 |
an:2024 |
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs |
An K, Chen Q, Deng C, Du Z, Gao C, Gao Z, Gu Y, He T, Hu H, Hu K, others (2024) FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs. arXiv preprint arXiv:2... |
docs/zh/part12/ch42_voice_style_control_dataset.md |
394 |
3 |
du:2024 |
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens |
Du Z, Chen Q, Zhang S, Hu K, Lu H, Yang Y, Hu H, Zheng S, Gu Y, Ma Z, Gao Z, Yan Z (2024) CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens. arXiv preprint arX... |
docs/zh/part12/ch42_voice_style_control_dataset.md |
396 |
4 |
du:2024 |
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models |
Du Z, Wang Y, Chen Q, Shi X, Lv X, Zhao T, Gao Z, Yang Y, Gao C, Wang H, others (2024) CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models. arXiv preprint arXiv:2412.10117. |
docs/zh/part12/ch42_voice_style_control_dataset.md |
398 |
5 |
mittag:2021 |
NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets |
Mittag G, Naderi B, Chehadi A, Möller S (2021) NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets. In: Interspeech 2021, pp 2127-2131. |
docs/zh/part12/ch42_voice_style_control_dataset.md |
402 |
7 |
yang:2025 |
Qwen3 Technical Report |
Yang A, Li A, Yang B, Zhang B, Hui B, Zheng B, Yu B, Gao C, Huang C, Lv C, others (2025) Qwen3 Technical Report. arXiv preprint arXiv:2505.09388. |
docs/zh/part12/ch43_latent_switch_69k.md |
287 |
1 |
wei:2022 |
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models |
1. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. |
docs/zh/part12/ch43_latent_switch_69k.md |
288 |
2 |
lightman:2023 |
Let's Verify Step by Step |
2. Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2023). Let's Verify Step by Step. arXiv:2305.20050. |
docs/zh/part12/ch43_latent_switch_69k.md |
289 |
3 |
yao:2023 |
ReAct: Synergizing Reasoning and Acting in Language Models |
3. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629. |
docs/zh/part12/ch43_latent_switch_69k.md |
290 |
4 |
deepseekai:2025 |
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning |
4. DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. |
docs/zh/part12/ch43_latent_switch_69k.md |
291 |
5 |
hendrycks:2021 |
Measuring Mathematical Problem Solving With the MATH Dataset |
5. Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., & Steinhardt, J. (2021). Measuring Mathematical Problem Solving With the MATH Dataset. NeurIPS 2021. |
docs/zh/part13/ch44_pretrain_recipes.md |
275 |
1 |
bavarian:2022 |
Efficient Training of Language Models to Fill in the Middle (FIM) |
Bavarian M, Jun H, Tezak N, Schulman J, McLeavey C, Tworek J, Chen M (2022) Efficient Training of Language Models to Fill in the Middle (FIM). arXiv preprint arXiv:2207.14255. |
docs/zh/part13/ch44_pretrain_recipes.md |
279 |
3 |
broder:1997 |
On the Resemblance and Containment of Documents |
Broder A Z (1997) On the Resemblance and Containment of Documents. In: Proceedings of the Compression and Complexity of Sequences, pp 21-29. |
docs/zh/part13/ch44_pretrain_recipes.md |
285 |
6 |
hoffmann:2022 |
Training Compute-Optimal Large Language Models (Chinchilla) |
Hoffmann J, Borgeaud S, Mensch A, Buchatskaya E, Cai T, Rutherford E, de Las Casas D, Hendricks L A, Welbl J, Clark A, others (2022) Training Compute-Optimal Large Language Models (Chinchilla). arXiv preprint arXiv:22... |
docs/zh/part13/ch44_pretrain_recipes.md |
295 |
11 |
sennrich:2016 |
Neural Machine Translation of Rare Words with Subword Units (BPE) |
Sennrich R, Haddow B, Birch A (2016) Neural Machine Translation of Rare Words with Subword Units (BPE). In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pp 1715-1725. |
docs/zh/part13/ch45_posttrain_recipes.md |
357 |
1 |
wang:2023 |
Self-Instruct: Aligning Language Models with Self-Generated Instructions |
Wang Y, Kordi Y, Mishra S, Liu A, Smith N A, Khashabi D, Hajishirzi H (2023) Self-Instruct: Aligning Language Models with Self-Generated Instructions. Proceedings of the 61st Annual Meeting of the Association for Comp... |
docs/zh/part13/ch45_posttrain_recipes.md |
363 |
4 |
ethayarajh:2024 |
Model Alignment as Prospect Theoretic Optimization |
Ethayarajh K, Xu W, Muennighoff N, Jurafsky D, Kiela D (2024) Model Alignment as Prospect Theoretic Optimization. Proceedings of the 41st International Conference on Machine Learning, pp 12634-12651. |
docs/zh/part13/ch45_posttrain_recipes.md |
365 |
5 |
gheshlaghi:2024 |
A General Theoretical Paradigm to Understand Learning from Human Preferences |
Gheshlaghi Azar M, Guo Z D, Piot B, Munos R, Rowland M, Valko M, Calandriello D (2024) A General Theoretical Paradigm to Understand Learning from Human Preferences. Proceedings of the 27th International Conference on ... |
docs/zh/part13/ch45_posttrain_recipes.md |
367 |
6 |
grattafiori:2024 |
The Llama 3 Herd of Models |
Grattafiori A, Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, Letman A, Mathur A, Schelten A, Vaughan A, others (2024) The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783. |
docs/zh/part13/ch45_posttrain_recipes.md |
369 |
7 |
lambert:2025 |
Tülu 3: Pushing Frontiers in Open Language Model Post-Training |
Lambert N, Morrison J, Pyatkin V, Huang S, Ivison H, Brahman F, Miranda L J V, Liu A, Dziri N, Lyu X, Gu Y, Malik S, Graf V, Hwang J D, Yang J, Le Bras R, Tafjord O, Wilhelm C, Soldaini L, Smith N A, Wang Y, Dasigi P,... |
docs/zh/part13/ch45_posttrain_recipes.md |
371 |
8 |
yang:2025 |
Qwen3 Technical Report |
Yang A, Li A, Yang B, Zhang B, Hui B, Zheng B, Yu B, Gao C, Huang C, Lv C, others (2025) Qwen3 Technical Report. arXiv preprint arXiv:2505.09388. |
docs/zh/part13/ch45_posttrain_recipes.md |
377 |
11 |
xu:2025 |
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing |
Xu Z, Jiang F, Niu L, Deng Y, Poovendran R, Choi Y, Lin B Y (2025) Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing. International Conference on Learning Representations. |
docs/zh/part13/ch45_posttrain_recipes.md |
383 |
14 |
singhal:2024 |
A Long Way to Go: Investigating Length Correlations in RLHF |
Singhal P, Goyal T, Xu J, Durrett G (2024) A Long Way to Go: Investigating Length Correlations in RLHF. First Conference on Language Modeling. |
docs/zh/part13/ch45_posttrain_recipes.md |
385 |
15 |
zhou:2023 |
Don't Make Your LLM an Evaluation Benchmark Cheater |
Zhou K, Zhu Y, Chen Z, Chen W, Zhao W X, Chen X, Lin Y, Wen J-R, Han J (2023) Don't Make Your LLM an Evaluation Benchmark Cheater. arXiv preprint arXiv:2311.01964. |
docs/zh/part13/ch45_posttrain_recipes.md |
387 |
16 |
lightman:2024 |
Let's Verify Step by Step |
Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, Leike J, Schulman J, Sutskever I, Cobbe K (2024) Let's Verify Step by Step. International Conference on Learning Representations. |
docs/zh/part13/ch46_rl_reasoning_data.md |
544 |
2 |
team:2025 |
Kimi k1.5: Scaling Reinforcement Learning with LLMs |
Team Kimi, Du A, Gao B, Xing B, Jiang C, Chen C, Li C, Xiao C, Du C, Liao C, others (2025) Kimi k1.5: Scaling Reinforcement Learning with LLMs. arXiv preprint arXiv:2501.12599. |
docs/zh/part13/ch46_rl_reasoning_data.md |
546 |
3 |
touvron:2023 |
Llama 2: Open Foundation and Fine-Tuned Chat Models |
Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, Bashlykov N, Batra S, Bhargava P, Bhosale S, others (2023) Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:2307.09288. |
docs/zh/part13/ch46_rl_reasoning_data.md |
552 |
6 |
zhou:2023 |
LIMA: Less Is More for Alignment |
Zhou C, Liu P, Xu P, Iyer S, Sun J, Mao Y, Ma X, Efrat A, Yu P, Yu L, Zhang S, Ghosh G, Lewis M, Zettlemoyer L, Levy O (2023) LIMA: Less Is More for Alignment. Advances in Neural Information Processing Systems, 36, 55... |
docs/zh/part13/ch46_rl_reasoning_data.md |
554 |
7 |
zelikman:2022 |
STaR: Bootstrapping Reasoning with Reasoning |
Zelikman E, Wu Y, Mu J, Goodman N (2022) STaR: Bootstrapping Reasoning with Reasoning. Advances in Neural Information Processing Systems, 35, 15476-15488. |
docs/zh/part13/ch46_rl_reasoning_data.md |
556 |
8 |
madaan:2023 |
Self-Refine: Iterative Refinement with Self-Feedback |
Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, Alon U, Dziri N, Prabhumoye S, Yang Y, Gupta S, Majumder B P, Hermann K, Welleck S, Yazdanbakhsh A, Clark P (2023) Self-Refine: Iterative Refinement with Se... |
docs/zh/part13/ch46_rl_reasoning_data.md |
558 |
9 |
lightman:2024 |
Let's Verify Step by Step |
Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, Leike J, Schulman J, Sutskever I, Cobbe K (2024) Let's Verify Step by Step. International Conference on Learning Representations. |