View and convert an HTML file in .NET application

Blog category: Imaging; .NET

October 1, 2026

HTML file is a file containing the document markup in HyperText Markup Language (HTML) intended for displaying the document in a web browser. HTML is typically used in conjunction with Cascading Style Sheets (CSS) and a scripting language such as JavaScript.

Web browsers such as Google Chrome, Microsoft Edge, Yandex Browser, Opera, UC Browser, Baidu Browser and others use the Chromium engine to render HTML pages.

The Mozilla Firefox web browser uses its own Gecko engine to render HTML pages.

When VintaSoft needed an HTML rendering solution, we evaluated all available options and decided to build our own HTML rendering engine. We realize that our engine is unlikely to match the quality and speed of Chromium or Gecko anytime soon, but that is not our goal. We do not claim that our engine can perfectly render every possible HTML page; rather, our objective is to provide accurate, high-speed rendering of simple to moderately complex HTML pages for our clients. VintaSoft’s HTML rendering engine is written entirely in C#, ensuring maximum performance within the .NET environment and cross-platform support.


VintaSoft Imaging .NET SDK version 15.2 and later allows to work with HTML files, i.e. it is possible to:
To work with an HTML file, the following nuget-packages are required: Vintasoft.Shared, Vintasoft.Imaging.

Since version 15.2 of VintaSoft Imaging .NET SDK:


Here is the C# code that demonstrates how to extract text from an HTML file:
/// <summary>
/// Extracts the text from HTML.
/// </summary>
/// <param name="htmlFileName">Name of the HTML file.</param>
public static void ExtractTextFromHtml(string htmlFileName)
{
    using (Vintasoft.Imaging.ImageCollection images = new Vintasoft.Imaging.ImageCollection())
    {
        // open HTML document
        images.Add(htmlFileName);

        try
        {
            // page number
            int pageNumber = 1;
            
            // for each page of HTML dcoument
            foreach (Vintasoft.Imaging.VintasoftImage image in images)
            {
                // write page number
                System.Console.WriteLine("\tPage Number: {0}", pageNumber++);
                System.Console.WriteLine();
                System.Console.WriteLine();

                // find text region metadata
                Vintasoft.Imaging.Metadata.TextRegionMetadata textRegionMetadata =
                    image.Metadata.MetadataTree.FindChildNode<Vintasoft.Imaging.Metadata.TextRegionMetadata>();

                // if current page has text content
                if (textRegionMetadata != null)
                {
                    // get text region
                    Vintasoft.Imaging.Text.TextRegion textRegion = textRegionMetadata.GetTextRegion();
                    // write page text content
                    System.Console.Write(textRegion.TextContent);
                }

                // if page separator must be added between pages
                if (pageNumber < images.Count)
                {
                    System.Console.WriteLine();
                    System.Console.WriteLine();
                    System.Console.WriteLine();
                }
            }
        }
        finally
        {
            // clear and dispose images
            images.ClearAndDisposeItems();
        }
    }
}


Here is the C# code that demonstrates how to convert an HTML file to a TIFF file:
/// <summary>
/// Converts HTML document to TIFF file using ImageCollection and TiffEncoder classes.
/// </summary>
public static void ConvertHtmlToTiff_ImageCollection(string htmlFileName, string tiffFileName)
{
    // create image collection
    using (Vintasoft.Imaging.ImageCollection images = 
        new Vintasoft.Imaging.ImageCollection())
    {
        // layout HTML document on page area: width - 210mm,  height - infinite
        Vintasoft.Imaging.Codecs.Decoders.HtmlLayoutSettings layoutSettings =
            new Vintasoft.Imaging.Codecs.Decoders.HtmlLayoutSettings();
        layoutSettings.PageLayoutSettings.PageSize = 
            Vintasoft.Imaging.ImageSize.FromMillimeters(210, 0, new Vintasoft.Imaging.Resolution(96));
        images.LayoutSettings.SetSettings(layoutSettings);

        // add HTML document to collection
        images.Add(htmlFileName);

        // render HTML pages in 144 DPI
        Vintasoft.Imaging.Codecs.Decoders.RenderingSettings renderingSettings = 
            new Vintasoft.Imaging.Codecs.Decoders.RenderingSettings();
        renderingSettings.Resolution = new Vintasoft.Imaging.Resolution(144);
        images.SetRenderingSettings(renderingSettings);

        // create TiffEncoder
        using (Vintasoft.Imaging.Codecs.Encoders.TiffEncoder encoder = 
            new Vintasoft.Imaging.Codecs.Encoders.TiffEncoder(true))
        {
            // set TIFF compression to Jpeg
            encoder.Settings.Compression = 
                Vintasoft.Imaging.Codecs.ImageFiles.Tiff.TiffCompression.Jpeg;

            // save images of image collection to TIFF file using TiffEncoder
            images.SaveSync(tiffFileName, encoder);
        }

        // dispose images
        images.ClearAndDisposeItems();
    }
}


Here is the C# code that demonstrates how to convert an HTML file to a PDF file:
/// <summary>
/// Converts HTML document to PDF document using ImageCollection and PdfEncoder classes.
/// </summary>
public static void ConvertHtmlToPdf_ImageCollection(string htmlFileName, string pdfFileName)
{
    // create image collection
    using (Vintasoft.Imaging.ImageCollection images = 
        new Vintasoft.Imaging.ImageCollection())
    {
        // layout HTML document on page area: width - 210mm,  height - infinite
        Vintasoft.Imaging.Codecs.Decoders.HtmlLayoutSettings layoutSettings =
            new Vintasoft.Imaging.Codecs.Decoders.HtmlLayoutSettings();
        layoutSettings.PageLayoutSettings.PageSize = 
            Vintasoft.Imaging.ImageSize.FromMillimeters(210, 0, new Vintasoft.Imaging.Resolution(96));
        images.LayoutSettings.SetSettings(layoutSettings);

        // add HTML document to collection
        images.Add(htmlFileName);
        
        // create PdfEncoder
        using (Vintasoft.Imaging.Codecs.Encoders.PdfEncoder encoder = 
            new Vintasoft.Imaging.Codecs.Encoders.PdfEncoder(true))
        {
            // set PDF resources compression to Jpeg
            encoder.Settings.Compression = Vintasoft.Imaging.Codecs.Encoders.PdfImageCompression.Jpeg;

            // save images of image collection to PDF document using PdfEncoder
            images.SaveSync(pdfFileName, encoder);
        }

        // dispose images
        images.ClearAndDisposeItems();
    }
}


Here is the C# code that demonstrates how to convert an HTML file to a DOCX file:
/// <summary>
/// Converts HTML document to a DOCX document.
/// </summary>
/// <param name="htmlFilePath">Path to a source HTML file.</param>
/// <param name="docxFilePath">Path to a result DOCX file.</param>
public static void ConvertHtmlToDocx(string htmlFilePath, string docxFilePath)
{
    Vintasoft.Imaging.Office.OpenXml.OpenXmlDocumentConverter.ConvertHtmlToDocx(htmlFilePath, docxFilePath);
}