E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL

stp2yJanuary 29, 20250 Comments

Architecture of OpenAI

[Submitted on 25 Sep 2024 (v1), last revised 28 Jan 2025 (this version, v2)]

View a PDF of the paper titled E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL, by Hasan Alp Caferou{g}lu and 1 other authors

View PDF
HTML (experimental)

Abstract:Translating Natural Language Queries into Structured Query Language (Text-to-SQL or NLQ-to-SQL) is a critical task extensively studied by both the natural language processing and database communities, aimed at providing a natural language interface to databases (NLIDB) and lowering the barrier for non-experts. Despite recent advancements made through the use of Large Language Models (LLMs), significant challenges remain. These include handling complex database schemas, resolving ambiguity in user queries, and generating SQL queries with intricate structures that accurately reflect the user’s intent. In this work, we introduce E-SQL, a novel pipeline specifically designed to address these challenges through direct schema linking and candidate predicate augmentation. E-SQL enhances the natural language query by incorporating relevant database items (i.e., tables, columns, and values) and conditions directly into the question and SQL construction plan, bridging the gap between the query and the database structure. The pipeline leverages candidate predicate augmentation to mitigate erroneous or incomplete predicates in generated SQLs. Comprehensive evaluations on the BIRD benchmark illustrate that E-SQL achieves competitive performance, particularly excelling in complex queries with a 66.29% execution accuracy on the test set. A further observation from our experiments reveals that incorporating schema filtering into the translation pipeline does not have a positive impact on performance when the most advanced proprietary LLMs are used. Additionally, our experiments with small LLMs highlight the importance and positive impact of enriched questions on their performance. Without fine-tuning, single-prompt SQL generation using enriched questions with DeepSeek Coder 7B Instruct 1.5v achieves 56.45% execution accuracy on the BIRD development set.

Submission history

From: Hasan Alp Caferoğlu [view email]
[v1]
Wed, 25 Sep 2024 09:02:48 UTC (729 KB)
[v2]
Tue, 28 Jan 2025 09:45:41 UTC (1,216 KB)

Source link
lol

By stp2y