Working with Compressed Resources in .NET
By Peter A. Bromberg, Ph.D.
Printer - Friendly Version
Peter Bromberg

The ability to use some of the built in methods of the Assembly class, such as GetManifestResourceStream, makes it easy for developers to embed resources of almost any type into an assembly and access these at runtime to be used in the functionality of the class itself.

You can embed almost any type of resource - images, Javascript files, text representations of XmlDocuments, and even delimited text files that have been exported from a database. But what about the idea of using Data Compression on some of these resources before they are embedded in your assembly? Based on some of the work I've done on compression, text files can often be compressed by as much as 90 percent. If you had the code to compress and decompress these resources, and the decompression code could be included in your assembly, then you would have the ingredients to access and decompress your in-assembly resources at runtime and since they would already be memory-resident at the time of decompression, you could potentially save a considerable amount of time.

In the case of an ASP.NET application, the one-time loading of the assembly could be stored in Application state, and would then be accessible to any ASP.NET page running in the application with no loading delays. And in the case of other applications, the rehydrated resource could be stored in the appDomain cache instead, basically accomplishing the same effect.

As it turns out, Mike Kruger's SharpZipLib C# codec provides the source code we need, and we would only need to include those classes necessary to decompress an embedded resource in order to meet our needs, so the impact of the decompression code on the size of our final assembly is minimal.

As a "Proof of Concept" exercise, I took an export of a database we have of famous quotations; said export consisting of 8,627 records each containing an Author, Author Info, Subject and a Quotation column, and put this into a delimited CSV text file which takes up 1186 Kb on the hard drive. After applying standard compression with a maximum compression factor from SharpZipLib, the resultant resource is only 423K - a 64% space savings. Here are a couple of the uncompressed records so that you can see what I started with:

Meister Eckhart|1260-1326 AD, German Mystic|Compassion|You may call God love, you may call God goodness. But the best name for God is compassion.
Dante (Alighieri)|1265-1321, Italian Philosopher, Poet|Conscience|O conscience, upright and stainless, how bitter a sting to thee is a little fault!

Along with all the code to decompress this embedded resource into a DataTable and provide two separate DataTable search methods FindQuotesByAuthor, and FindQuotesBySubject, each using the Select filter method on the rehydrated DataTable in memory, the final QuotationControl assembly is just 472K - well within the range of acceptable DLL sizes to be loaded into memory one time upon Application instantiation.

The Select method is very fast, even though the DataTable is not indexed, because the records are all in-memory and so there is no disk I/O overhead as with a database. I've left out primary keys and indexing here in order to keep the focus on the basic premise, which is that of being able to decompress an embedded resource at runtime.

However, though ADO.NET does not expose indexes, it does create and use them both for Primary Keys and for DataViews. You can take advantage of the indexes created for Primary Keys or DataViews as follows:

If you are searching on a Primary Key field, or if you can make the field(s) you are searching on a Primary Key, then instead of DataTable.Select(), you can use DataTable.Rows.Find(). Find() will look up a primary key value using the built index.

If you are performing repeated searches on a particular column or set of columns, you can create a DataView sorted on those columns, and then use DataView.FindRows() to search for a value within the columns. DataViews create an index for the sorted columns, and FindRows() uses this. Also, when using a DataView, it is more efficient to use the constructor that accepts the table, filter, sort, and rowversions all at once. Each time you set the filter, sort, or rowversion properties the index is recreated, so creating a DataView and setting these properties individually actually creates the index several times. An example is shown immediately below:

private void MakeDataView(DataSet ds)
{
DataView dv = new DataView(ds.Tables["Suppliers"], "Country = 'UK'", "CompanyName", DataViewRowState.CurrentRows);
dv.AllowEdit = true;
dv.AllowNew = true;
dv.AllowDelete = true;
}

I've included two solutions in the downloadable source code for this article:

The first solution is called NZipLibForm and is just an updated test harness for SharpZipLib that allows you to choose Compress, get a file dialog, find the resource you want compressed, and have it automatically saved in the same folder with the extension ".dat" appended to it. By leaving the "Compress to/from file" checkbox unchecked, you can also paste a resource into the upper text box and see the compressed result in the lower textbox.

The second solution is the QuotationControl (its not really a control, just a class library, but that's the way I started it out so I've been reluctant to rename it). This is the real meat of the madness, and basically looks like this:

using System;
using System.Web.UI;
using System.Reflection;
using System.IO;
using System.Data;
using System.Diagnostics;
using ICSharpCode.SharpZipLib;
using System.Collections;

namespace QuotationControl
{
public class QuotationLookup
{
private string theQuotes=String.Empty;
public DataTable QuotesTable = new DataTable();
private DataTable foundQuotes=new DataTable();
public DataTable FoundQuotes
{
get
{
return foundQuotes;
}
}

public void FindQuotesByAuthor (string Author)
{
foundQuotes.Clear();
string strExpr;
strExpr = "Author LIKE '%" +Author.Trim() + "%'";
// Use the Select method to find all rows matching the filter.
try
{
DataRow[] foundRows =
QuotesTable.Select(strExpr);

for(int j =0;j<foundRows.Length;j++)
{
object[] theRow={foundRows[j].ItemArray[0].ToString(),
              foundRows[j].ItemArray[1].ToString(),
              foundRows[j].ItemArray[2].ToString(),
              foundRows[j].ItemArray[3].ToString()};
foundQuotes.Rows.Add(theRow);
}

foundQuotes.AcceptChanges();
}
catch(Exception ex)
{throw new Exception(ex.Message);
}
}


public void FindQuotesBySubject (string Subject)
{
foundQuotes.Clear();
string strExpr;
strExpr = "Subject LIKE '%" +Subject.Trim() + "%'";
try
{
DataRow[] foundRows =
QuotesTable.Select(strExpr);

for(int j =0;j<foundRows.Length;j++)
{
object[] theRow={foundRows[j].ItemArray[0].ToString(),
                 foundRows[j].ItemArray[1].ToString(),
                 foundRows[j].ItemArray[2].ToString(),
                 foundRows[j].ItemArray[3].ToString()}; 
foundQuotes.Rows.Add(theRow);
}
foundQuotes.AcceptChanges();
}
catch(Exception ex)
{
throw new Exception(ex.Message);
}
}

public QuotationLookup()
{
foundQuotes.Columns.Add("Author");
foundQuotes.Columns.Add("AuthorInfo");
foundQuotes.Columns.Add("Subject");
foundQuotes.Columns.Add("Quotation");

theQuotes = GetDecompressedResourceString("QuotationControl.quotations.dat");
QuotesTable.Columns.Add( "Author", typeof(string) );
QuotesTable.Columns.Add( "AuthorInfo", typeof(string) );
QuotesTable.Columns.Add( "Subject", typeof(string) );
QuotesTable.Columns.Add( "Quotation", typeof(string) );
theQuotes=theQuotes.Replace("\"","");
string[] whizzies = theQuotes.Split(new Char[] {'\n'});
for (int i=0;i<whizzies.Length;i++)
{
object[] theRow=whizzies[i].Split(new Char[] {'|'});
QuotesTable.Rows.Add( theRow);
}
QuotesTable.AcceptChanges();
}

private string GetDecompressedResourceString(string resource)
{
try
{
Assembly asm = Assembly.GetExecutingAssembly();
Debug.WriteLine(asm.FullName);
Stream stm=asm.GetManifestResourceStream(resource);
Debug.WriteLine(stm.Length.ToString());
BinaryReader br = new BinaryReader(stm);
long siz = stm.Length;
byte[] bytInput =null;
bytInput=br.ReadBytes((int)siz);
string theResource=DeCompress(bytInput);
br.Close();
stm.Close();
return theResource;
}
catch(Exception ex)
{
throw new ApplicationException(ex.Message);
}
}

private string DeCompress(byte[] bytInput)
{
string strResult="";
int totalLength = 0;
byte[] writeData = new byte[4096];
Stream s2 = new ICSharpCode.SharpZipLib.Zip.Compression.Streams.InflaterInputStream(new MemoryStream(bytInput));
try
{
while (true)
{
int size = s2.Read(writeData, 0, writeData.Length);
if (size > 0)
{
totalLength += size;
strResult+=System.Text.Encoding.UTF8.GetString(writeData, 0,
size);
}
else
{
break;
}
}
s2.Close();
return strResult;
}
catch(Exception e)
{
throw new Exception(e.ToString());

}
}
}
}

Note that I've left out the SharpZipLib code classes in the above snippet, but basically the DeCompress method encapsulates all the code necessary to invoke the decompression engine on your passed in resource stream. Note also that it is necessary to use the BinaryReader to get your resource out of the assembly in the correct format for decompression to work!

Once the resource is decompressed into a string, it is a relatively straightforward process to use the Split method, once on the newline character to get all your rows into an array, and then again on the pipe ("|") character which is what I chose as my column delimiter when exporting these records to a text file from Sql Server, in order to get the column items for each row in the DataTable.

You can apply this overall technique to tabular data such as the above (you could also construct an entire DataSet in memory complete with indexes, primary keys, and constraints), it can be applied to a large XmlDocument which is then rehydrated and stored in - memory, and so on.

If you'd like to try out the result before downloading the sample solutions, just hit this demo page. For an Author search, you might try the search term "shak" (for Shakespeare) or "Bacon" (for Francis Bacon), and for a subject search, why not try "Happiness" - a good subject for all of us! Interestingly, as a philosophical side note, since my search term is surrounded by "%" wildcards, if you search on "war" you will also be presented with the quotations on the subjects of "cowardice" and "rewards". Enjoy.

Download the Code that accompanies this article


Peter Bromberg is a C# MVP, MCP, and .NET consultant who has worked in the banking and financial industry for 20 years. He has architected and developed web - based corporate distributed application solutions since 1995, and focuses exclusively on the .NET Platform.