References
[1] Goodside, R. (2022). Exploiting GPT-3 prompts with malicious inputs that order the model to ignore its previous directions.
[2] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173.
[3] Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., & Wang, K. (2023). Prompt Injection attack against LLM-integrated Applications. arXiv:2306.05499.
[4] Biggio, B., Nelson, B., & Laskov, P. (2012). Poisoning attacks against support vector machines. ICML 2012.
[5] Wan, A., Wallace, E., Shen, S., & Klein, D. (2023). Poisoning Language Models During Instruction Tuning. arXiv:2302.07815.
[6] OWASP. (2023). OWASP Top 10 for Large Language Model Applications.
[7] Deng, G., Liu, Y., Li, Y., Wang, K., Zhang, Y., & Li, Z. (2023). MasterKey: Automated Jailbreaking of Large Language Model Chatbots. NDSS 2024.
[8] 國家新一代人工智能治理專業委員會. 《新一代人工智能治理原則——發展負責任的人工智能》. 2019年6月.
[9] 中華人民共和國國務院. 《新一代人工智能發展規劃》(國發〔2017〕35號). 2017年7月.