Detect tables in a document image in .NET application

Blog category: Imaging; .NET

October 1, 2026

Documents often contain tabular data - information organized into a table with rows and columns.

If all cells in a table are enclosed by lines, the table is considered to have borders. Conversely, if most cells lack enclosing lines, it is a borderless table.

Recognition of tables with borders in a document image can be implemented algorithmically, because to complete the task, it is necessary to recognize lines in the document image and find the lines, which define tables in the document.

Recognition of borderless tables in a document image is a more complex task because it requires identifying blocks of information in the image and recognizing that these blocks of information are presented as tabular data. Nowadays, this task is typically accomplished using artificial intelligence.


VintaSoft Imaging .NET SDK 15.2 implements recognition of tables with borders, i.e. SDK has algorithm that can recognize boundaries of the table, table rows, table cells in a document image. The algorithm yields good results if the document image is of high quality - specifically, with a resolution of 300 dpi or higher.


The SDK also allows to utilize table recognition results for practical purposes:


Here is C# code that demonstrates how to convert a raster image into a DOCX document (including text recognition and table detection):
using Vintasoft.Imaging;
using Vintasoft.Imaging.ImageProcessing.Document;
using Vintasoft.Imaging.ImageProcessing.Info.TableDetection;
using Vintasoft.Imaging.Ocr;
using Vintasoft.Imaging.Ocr.Results;
using Vintasoft.Imaging.Ocr.Tesseract;
using Vintasoft.Imaging.Pdf;
using Vintasoft.Imaging.Pdf.Ocr;
using Vintasoft.Imaging.Pdf.Office;

class ConvertRasterImagesToDocxDocumentExample
{
    // Required assemblies to run this code:
    // Vintasoft.Shared.dll, Vintasoft.Imaging.dll,
    // Vintasoft.Imaging.DocCleanup.dll,
    // Vintasoft.Imaging.Office.OpenXml.dll,
    // Vintasoft.Imaging.Ocr.dll, Vintasoft.Imaging.Ocr.Tesseract.dll,
    // Vintasoft.Imaging.Pdf.dll, Vintasoft.Imaging.Pdf.Office.dll, Vintasoft.Imaging.Pdf.Ocr.dll

    /// <summary>
    /// Converts a raster image file to a DOCX document (with OCR and table detection).
    /// </summary>
    /// <param name="ocrLanguage">An OCR language.</param>
    /// <param name="imageFilename">A filename of source raster image (TIFF, PNG, ...) file.</param>
    /// <param name="docxFilename">A filename of destination DOCX file.</param>
    public static void ConvertRasterImagesToDocxDocument(
        OcrLanguage ocrLanguage, string imageFilename, string docxFilename)
    {
        // create an image collection
        using (ImageCollection images = new ImageCollection())
        {
            // add images from image file into image collection
            images.Add(imageFilename);

            // create a searchable PDF document
            using (PdfDocument document = new PdfDocument())
            {
                // create a PDF document builder
                PdfDocumentBuilder documentBuilder = new PdfDocumentBuilder(document);

                // specify that text must be placed over image
                documentBuilder.PageCreationMode = PdfPageCreationMode.TextOverImage;

                // create the Tesseract OCR engine
                using (TesseractOcr tesseractOcr = new TesseractOcr(@".\TesseractOCR"))
                {
                    // create OCR settings
                    OcrEngineSettings ocrSettings = new OcrEngineSettings(ocrLanguage);

                    // create the OCR engine manager
                    OcrEngineManager engineManager = new OcrEngineManager(tesseractOcr);

                    // OCR Preprocessing:
                    // AutoInvert, HalftoneRemoval, BorderClear, Deskew, HolePunchRemoval,
                    // Despeckle, AutoTextOrientation, Segmentation, TableWithBordersDetection
                    OcrPreprocessingCommand ocrPreprocessing = new OcrPreprocessingCommand();
                    // disable binarization
                    ocrPreprocessing.Binarization = null;

                    // for each image in image collection
                    foreach (VintasoftImage image in images)
                    {
                        // execute preprocessing
                        ocrPreprocessing.ExecuteInPlace(image);

                        // recognize text on image
                        OcrPage page = engineManager.Recognize(image, ocrSettings, ocrPreprocessing.DetectedRegions);

                        // add recognized OCR page to the PDF document
                        documentBuilder.AddPage(image, page);
                    }

                    // shutdown OCR engine
                    tesseractOcr.Shutdown();

                    // clear and dispose images in image collection
                    images.ClearAndDisposeItems();

                    // create converter that allows to convert PDF document to a DOCX document
                    using (PdfToDocxConverter converter = new PdfToDocxConverter())
                    {
                        // set the destination DOCX file
                        converter.OutputFilename = docxFilename;

                        // set converter settings
                        converter.ConvertGraphics = true;
                        converter.DetectHeaderFooter = true;

                        // enable table detection
                        converter.DetectTables = true;
                        converter.TableDetectionCommand = new TableWithBordersDetectionCommand();

                        // convert searchable PDF document to a DOCX document
                        converter.Execute(document);
                    }
                }
            }
        }
    }
}