iceberg

package module
v0.6.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 21, 2026 License: Apache-2.0 Imports: 40 Imported by: 21

README

Iceberg Golang

Go Reference

iceberg is a Golang implementation of the Iceberg table spec.

Build From Source

Prerequisites
  • Go 1.25 or later
Build
$ git clone https://github.com/apache/iceberg-go.git
$ cd iceberg-go/cmd/iceberg && go build .

Running Tests

Use the Makefile so commands stay in sync with CI (e.g. golangci-lint version).

Unit tests
make test
Linting
make lint

Install the linter first

make lint-install
# or: go install github.com/golangci/golangci-lint/cmd/golangci-lint@v2.8.0
Integration tests

Prerequisites: Docker, Docker Compose

  1. Start the Docker containers using docker compose:

    make integration-setup
    
  2. Export the required environment variables:

    export AWS_S3_ENDPOINT=http://$(docker inspect -f '{{range.NetworkSettings.Networks}}{{.IPAddress}}{{end}}' minio):9000
    export AWS_REGION=us-east-1
    export SPARK_CONTAINER_ID=$(docker ps -qf 'name=spark-iceberg')
    export DOCKER_API_VER=$(docker version -f '{{.Server.APIVersion}}')
    
  3. Run the integration tests:

    make integration-test
    

    Or run a single suite: make integration-scanner, make integration-io, make integration-rest, make integration-spark.

Feature Support / Roadmap

FileSystem Support
Filesystem Type Supported
S3 X
Google Cloud Storage X
Azure Blob Storage X
Local Filesystem X
Metadata
Operation Supported
Get Schema X
Get Snapshots X
Get Sort Orders X
Get Partition Specs X
Get Manifests X
Create New Manifests X
Plan Scan x
Plan Scan for Snapshot x
Catalog Support
Operation REST Hive Glue SQL Hadoop
Load Table X X X X X
List Tables X X X X X
Create Table X X X X X
Register Table X X X
Update Current Snapshot X X X X X
Create New Snapshot X X X X X
Rename Table X X X X
Drop Table X X X X X
Alter Table X X X X X
Check Table Exists X X X X X
Set Table Properties X X X X X
List Namespaces X X X X X
Create Namespace X X X X X
Check Namespace Exists X X X X X
Drop Namespace X X X X X
Update Namespace Properties X X X X
Create View X X X
Load View X X
List View X X X
Drop View X X X
Check View Exists X X X
Read/Write Data Support
  • Data can currently be read as an Arrow Table or as a stream of Arrow record batches.
Supported Write Operations

As long as the FileSystem is supported and the Catalog supports altering the table, the following tracks the current write support:

Operation Supported
Append Stream X
Append Data Files X
Rewrite Files
Rewrite manifests
Overwrite Files X
Copy-On-Write Delete X
Write Pos Delete X
Write Eq Delete
Row Delta
CLI Usage

Run go build ./cmd/iceberg from the root of this repository to build the CLI executable, alternately you can run go install github.com/apache/iceberg-go/cmd/iceberg@latest to install it to the bin directory of your GOPATH.

The iceberg CLI usage is very similar to pyiceberg CLI
You can pass the catalog URI with --uri argument.

Example: You can start the Iceberg REST API docker image which runs on default in port 8181

docker pull apache/iceberg-rest-fixture:latest
docker run -p 8181:8181 apache/iceberg-rest-fixture:latest

and run the iceberg CLI pointing to the REST API server.

 ./iceberg --uri http://0.0.0.0:8181 list
┌─────┐
| IDs |
| --- |
└─────┘

Create Namespace

./iceberg --uri http://0.0.0.0:8181 create namespace taxitrips

List Namespace

 ./iceberg --uri http://0.0.0.0:8181 list
┌───────────┐
| IDs       |
| --------- |
| taxitrips |
└───────────┘


Get in Touch

Documentation

Index

Constants

View Source
const (
	// RowIDFieldID is the field ID for _row_id (optional long). A unique long identifier for every row.
	RowIDFieldID = 2147483540
	// LastUpdatedSequenceNumberFieldID is the field ID for _last_updated_sequence_number (optional long).
	// The sequence number of the commit that last updated the row.
	LastUpdatedSequenceNumberFieldID = 2147483539
)

Row lineage metadata column field IDs (v3+). Reserved IDs are Integer.MAX_VALUE - 107 and 108 per the Iceberg spec (Metadata Columns / Row Lineage).

View Source
const (
	RowIDColumnName                     = "_row_id"
	LastUpdatedSequenceNumberColumnName = "_last_updated_sequence_number"
)

Row lineage metadata column names (v3+).

View Source
const (
	PartitionDataIDStart   = 1000
	InitialPartitionSpecID = 0
)

Variables

View Source
var (
	ErrInvalidTypeString       = errors.New("invalid type")
	ErrNotImplemented          = errors.New("not implemented")
	ErrInvalidArgument         = errors.New("invalid argument")
	ErrInvalidFormatVersion    = fmt.Errorf("%w: invalid format version", ErrInvalidArgument)
	ErrInvalidSchema           = errors.New("invalid schema")
	ErrInvalidPartitionSpec    = errors.New("invalid partition spec")
	ErrInvalidTransform        = errors.New("invalid transform syntax")
	ErrType                    = errors.New("type error")
	ErrBadCast                 = errors.New("could not cast value")
	ErrBadLiteral              = errors.New("invalid literal value")
	ErrInvalidBinSerialization = errors.New("invalid binary serialization")
	ErrResolve                 = errors.New("cannot resolve type")
)
View Source
var PositionalDeleteSchema = NewSchema(0,
	NestedField{ID: 2147483546, Type: PrimitiveTypes.String, Name: "file_path", Required: true},
	NestedField{ID: 2147483545, Type: PrimitiveTypes.Int64, Name: "pos", Required: true},
)
View Source
var PrimitiveTypes = struct {
	Bool          PrimitiveType
	Int32         PrimitiveType
	Int64         PrimitiveType
	Float32       PrimitiveType
	Float64       PrimitiveType
	Date          PrimitiveType
	Time          PrimitiveType
	Timestamp     PrimitiveType
	TimestampTz   PrimitiveType
	TimestampNs   PrimitiveType
	TimestampTzNs PrimitiveType
	String        PrimitiveType
	Binary        PrimitiveType
	UUID          PrimitiveType
	Unknown       PrimitiveType
}{
	Bool:          BooleanType{},
	Int32:         Int32Type{},
	Int64:         Int64Type{},
	Float32:       Float32Type{},
	Float64:       Float64Type{},
	Date:          DateType{},
	Time:          TimeType{},
	Timestamp:     TimestampType{},
	TimestampTz:   TimestampTzType{},
	TimestampNs:   TimestampNsType{},
	TimestampTzNs: TimestampTzNsType{},
	String:        StringType{},
	Binary:        BinaryType{},
	UUID:          UUIDType{},
	Unknown:       UnknownType{},
}
View Source
var UnpartitionedSpec = &PartitionSpec{id: 0}

UnpartitionedSpec is the default unpartitioned spec which can be used for comparisons or to just provide a convenience for referencing the same unpartitioned spec object.

Functions

func ExpressionEvaluator

func ExpressionEvaluator(s *Schema, unbound BooleanExpression, caseSensitive bool) (func(StructLike) (bool, error), error)

ExpressionEvaluator returns a function which can be used to evaluate a given expression as long as a structlike value is passed which operates like and matches the passed in schema.

func ExtractFieldIDs

func ExtractFieldIDs(expr BooleanExpression) ([]int, error)

ExtractFieldIDs returns a slice containing the field IDs which are referenced by any terms in the given expression. This enables retrieving exactly which fields are needed for an expression.

func GeneratePartitionFieldName added in v0.4.0

func GeneratePartitionFieldName(schema *Schema, field PartitionField) (string, error)

GeneratePartitionFieldName returns default partition field name based on field transform type

The default names are aligned with other client implementations https://github.com/apache/iceberg/blob/main/core/src/main/java/org/apache/iceberg/BaseUpdatePartitionSpec.java#L518-L563

func IndexByID

func IndexByID(schema *Schema) (map[int]NestedField, error)

IndexByID performs a post-order traversal of the given schema and returns a mapping from field ID to field.

func IndexByName

func IndexByName(schema *Schema) (map[string]int, error)

IndexByName performs a post-order traversal of the schema and returns a mapping from field name to field ID.

func IndexNameByID

func IndexNameByID(schema *Schema) (map[int]string, error)

IndexNameByID performs a post-order traversal of the schema and returns a mapping from field ID to field name.

func IndexParents

func IndexParents(schema *Schema) (map[int]int, error)

IndexParents generates an index of field IDs to their parent field IDs. Root fields are not indexed

func IsMetadataColumn added in v0.6.0

func IsMetadataColumn(fieldID int) bool

IsMetadataColumn returns true if the field ID is a reserved metadata column (e.g. row lineage).

func PreOrderVisit added in v0.2.0

func PreOrderVisit[T any](sc *Schema, visitor PreOrderSchemaVisitor[T]) (res T, err error)

func PropUInt added in v0.6.0

func PropUInt(p Properties, key string, defVal uint) uint

PropUInt reads an unsigned-integer property by key. A missing key, an unparseable value, or a negative value returns defVal — PropUInt uses strconv.ParseUint, which rejects negatives rather than silently wrapping them to a large positive number.

func SetSchemaCacheSize added in v0.6.0

func SetSchemaCacheSize(size int) error

SetSchemaCacheSize resizes the manifest-entry schema cache used by the DataFile codec. The default capacity is sized for a few thousand active partition specs; long-running consumers with larger working sets (e.g. a compaction service touching many tables) should raise it. Existing entries are preserved on grow; on shrink, least-recently used entries are evicted down to the new size. Not safe to call concurrently with codec operations.

func Version

func Version() string

func Visit

func Visit[T any](sc *Schema, visitor SchemaVisitor[T]) (res T, err error)

Visit accepts a visitor and performs a post-order traversal of the given schema.

func VisitBoundPredicate

func VisitBoundPredicate[T any](e BoundPredicate, visitor BoundBooleanExprVisitor[T]) T

VisitBoundPredicate uses a BoundBooleanExprVisitor to call the appropriate method based on the type of operation in the predicate. This is a convenience function for implementing the VisitBound method of a BoundBooleanExprVisitor by simply calling iceberg.VisitBoundPredicate(pred, this).

func VisitExpr

func VisitExpr[T any](expr BooleanExpression, visitor BooleanExprVisitor[T]) (res T, err error)

VisitExpr is a convenience function to use a given visitor to visit all parts of a boolean expression in-order. Values returned from the methods are passed to the subsequent methods, effectively "bubbling up" the results.

func VisitMappedFields added in v0.2.0

func VisitMappedFields[S, T any](fields []MappedField, visitor NameMappingVisitor[S, T]) (res S, err error)

func VisitNameMapping added in v0.2.0

func VisitNameMapping[S, T any](obj NameMapping, visitor NameMappingVisitor[S, T]) (res S, err error)

func VisitSchemaWithPartner

func VisitSchemaWithPartner[T, P any](sc *Schema, partner P, visitor SchemaWithPartnerVisitor[T, P], accessor PartnerAccessor[P]) (res T, err error)

func WriteManifestList

func WriteManifestList(version int, out io.Writer, snapshotID int64, parentSnapshotID, sequenceNumber *int64, firstRowId int64, files []ManifestFile) (err error)

WriteManifestList writes a list of manifest files to an avro file.

Types

type AboveMaxLiteral

type AboveMaxLiteral interface {
	Literal
	// contains filtered or unexported methods
}

AboveMaxLiteral represents values that are above the maximum for their type such as values > math.MaxInt32 for an Int32Literal

type AfterFieldVisitor

type AfterFieldVisitor interface {
	AfterField(field NestedField)
}

type AfterListElementVisitor

type AfterListElementVisitor interface {
	AfterListElement(elem NestedField)
}

type AfterMapKeyVisitor

type AfterMapKeyVisitor interface {
	AfterMapKey(key NestedField)
}

type AfterMapValueVisitor

type AfterMapValueVisitor interface {
	AfterMapValue(value NestedField)
}

type AlwaysFalse

type AlwaysFalse struct{}

AlwaysFalse is the boolean expression "False"

func (AlwaysFalse) Equals

func (AlwaysFalse) Equals(other BooleanExpression) bool

func (AlwaysFalse) Negate

func (AlwaysFalse) Negate() BooleanExpression

func (AlwaysFalse) Op

func (AlwaysFalse) Op() Operation

func (AlwaysFalse) String

func (AlwaysFalse) String() string

type AlwaysTrue

type AlwaysTrue struct{}

AlwaysTrue is the boolean expression "True"

func (AlwaysTrue) Equals

func (AlwaysTrue) Equals(other BooleanExpression) bool

func (AlwaysTrue) Negate

func (AlwaysTrue) Negate() BooleanExpression

func (AlwaysTrue) Op

func (AlwaysTrue) Op() Operation

func (AlwaysTrue) String

func (AlwaysTrue) String() string

type AndExpr

type AndExpr struct {
	// contains filtered or unexported fields
}

func (AndExpr) Equals

func (a AndExpr) Equals(other BooleanExpression) bool

func (AndExpr) Negate

func (a AndExpr) Negate() BooleanExpression

func (AndExpr) Op

func (AndExpr) Op() Operation

func (AndExpr) String

func (a AndExpr) String() string

type AvroEntryMarshaler added in v0.6.0

type AvroEntryMarshaler interface {
	MarshalAvroEntry(spec PartitionSpec, schema *Schema, version int) ([]byte, error)
}

AvroEntryMarshaler is implemented by DataFile values that can be encoded using the manifest-entry Avro encoding. The iceberg package's built-in DataFile implementation satisfies it; external implementations can also satisfy it to participate in the github.com/apache/iceberg-go/codec DataFile codec.

The encoded bytes are the same bytes a manifest carries for this data file. Implementations must produce output that the iceberg package's manifest-entry Avro decoder accepts.

type BeforeFieldVisitor

type BeforeFieldVisitor interface {
	BeforeField(field NestedField)
}

type BeforeListElementVisitor

type BeforeListElementVisitor interface {
	BeforeListElement(elem NestedField)
}

type BeforeMapKeyVisitor

type BeforeMapKeyVisitor interface {
	BeforeMapKey(key NestedField)
}

type BeforeMapValueVisitor

type BeforeMapValueVisitor interface {
	BeforeMapValue(value NestedField)
}

type BelowMinLiteral

type BelowMinLiteral interface {
	Literal
	// contains filtered or unexported methods
}

BelowMinLiteral represents values that are below the minimum for their type such as values < math.MinInt32 for an Int32Literal

type BinaryLiteral

type BinaryLiteral []byte

func (BinaryLiteral) Any added in v0.2.0

func (b BinaryLiteral) Any() any

func (BinaryLiteral) Comparator

func (BinaryLiteral) Comparator() Comparator[[]byte]

func (BinaryLiteral) Equals

func (b BinaryLiteral) Equals(other Literal) bool

func (BinaryLiteral) MarshalBinary

func (b BinaryLiteral) MarshalBinary() (data []byte, err error)

func (BinaryLiteral) String

func (b BinaryLiteral) String() string

func (BinaryLiteral) To

func (b BinaryLiteral) To(typ Type) (Literal, error)

func (BinaryLiteral) Type

func (b BinaryLiteral) Type() Type

func (*BinaryLiteral) UnmarshalBinary

func (b *BinaryLiteral) UnmarshalBinary(data []byte) error

func (BinaryLiteral) Value

func (b BinaryLiteral) Value() []byte

type BinaryType

type BinaryType struct{}

func (BinaryType) Equals

func (BinaryType) Equals(other Type) bool

func (BinaryType) String

func (BinaryType) String() string

func (BinaryType) Type

func (BinaryType) Type() string

type BoolLiteral

type BoolLiteral bool

func (BoolLiteral) Any added in v0.2.0

func (b BoolLiteral) Any() any

func (BoolLiteral) Comparator

func (BoolLiteral) Comparator() Comparator[bool]

func (BoolLiteral) Equals

func (b BoolLiteral) Equals(l Literal) bool

func (BoolLiteral) MarshalBinary

func (b BoolLiteral) MarshalBinary() (data []byte, err error)

func (BoolLiteral) String

func (b BoolLiteral) String() string

func (BoolLiteral) To

func (b BoolLiteral) To(t Type) (Literal, error)

func (BoolLiteral) Type

func (b BoolLiteral) Type() Type

func (*BoolLiteral) UnmarshalBinary

func (b *BoolLiteral) UnmarshalBinary(data []byte) error

func (BoolLiteral) Value

func (b BoolLiteral) Value() bool

type BooleanExprVisitor

type BooleanExprVisitor[T any] interface {
	VisitTrue() T
	VisitFalse() T
	VisitNot(childResult T) T
	VisitAnd(left, right T) T
	VisitOr(left, right T) T
	VisitUnbound(UnboundPredicate) T
	VisitBound(BoundPredicate) T
}

BooleanExprVisitor is an interface for recursively visiting the nodes of a boolean expression

type BooleanExpression

type BooleanExpression interface {
	fmt.Stringer
	Op() Operation
	Negate() BooleanExpression
	Equals(BooleanExpression) bool
}

BooleanExpression represents a full expression which will evaluate to a boolean value such as GreaterThan or StartsWith, etc.

func BindExpr

func BindExpr(s *Schema, expr BooleanExpression, caseSensitive bool) (BooleanExpression, error)

BindExpr recursively binds each portion of an expression using the provided schema. Because the expression can end up being simplified to just AlwaysTrue/AlwaysFalse, this returns a BooleanExpression.

func IsIn

func IsIn[T LiteralType](t UnboundTerm, vals ...T) BooleanExpression

IsIn is a convenience wrapper for constructing an unbound set predicate for OpIn. It returns a BooleanExpression instead of an UnboundPredicate because depending on the arguments, it can automatically reduce to AlwaysFalse or AlwaysTrue (if given no values for examples). It may also reduce to EqualTo if only one value is provided.

Will panic if t is nil

func NewAnd

func NewAnd(left, right BooleanExpression, addl ...BooleanExpression) BooleanExpression

NewAnd will construct a new AndExpr, allowing the caller to provide potentially more than just two arguments which will be folded to create an appropriate expression tree. i.e. NewAnd(a, b, c, d) becomes AndExpr(a, AndExpr(b, AndExpr(c, d)))

Slight optimizations are performed on creation if either argument is AlwaysFalse or AlwaysTrue by performing reductions. If any argument is AlwaysFalse, then everything will get folded to a return of AlwaysFalse. If an argument is AlwaysTrue, then the other argument will be returned directly rather than creating an AndExpr.

Will panic if any argument is nil

func NewNot

NewNot creates a BooleanExpression representing a "Not" operation on the given argument. It will optimize slightly though:

If the argument is AlwaysTrue or AlwaysFalse, the appropriate inverse expression will be returned directly. If the argument is itself a NotExpr, then the child will be returned rather than NotExpr(NotExpr(child)).

func NewOr

func NewOr(left, right BooleanExpression, addl ...BooleanExpression) BooleanExpression

NewOr will construct a new OrExpr, allowing the caller to provide potentially more than just two arguments which will be folded to create an appropriate expression tree. i.e. NewOr(a, b, c, d) becomes OrExpr(a, OrExpr(b, OrExpr(c, d)))

Slight optimizations are performed on creation if either argument is AlwaysFalse or AlwaysTrue by performing reductions. If any argument is AlwaysTrue, then everything will get folded to a return of AlwaysTrue. If an argument is AlwaysFalse, then the other argument will be returned directly rather than creating an OrExpr.

Will panic if any argument is nil

func NotIn

func NotIn[T LiteralType](t UnboundTerm, vals ...T) BooleanExpression

NotIn is a convenience wrapper for constructing an unbound set predicate for OpNotIn. It returns a BooleanExpression instead of an UnboundPredicate because depending on the arguments, it can automatically reduce to AlwaysFalse or AlwaysTrue (if given no values for examples). It may also reduce to NotEqualTo if only one value is provided.

Will panic if t is nil

func RewriteNotExpr

func RewriteNotExpr(expr BooleanExpression) (BooleanExpression, error)

RewriteNotExpr rewrites a boolean expression to remove "Not" nodes from the expression tree. This is because Projections assume there are no "not" nodes.

Not nodes will be replaced with simply calling `Negate` on the child in the tree.

func SetPredicate

func SetPredicate(op Operation, t UnboundTerm, lits []Literal) BooleanExpression

SetPredicate creates a boolean expression representing a predicate that uses a set of literals as the argument, like In or NotIn. Duplicate literals will be folded into a set, only maintaining the unique literals.

Will panic if op is not a valid Set operation

func TranslateColumnNames

func TranslateColumnNames(expr BooleanExpression, fileSchema *Schema) (BooleanExpression, error)

TranslateColumnNames converts the names of columns in an expression by looking up the field IDs in the file schema. If columns don't exist they are replaced with AlwaysFalse or AlwaysTrue depending on the operator.

type BooleanType

type BooleanType struct{}

func (BooleanType) Equals

func (BooleanType) Equals(other Type) bool

func (BooleanType) String

func (BooleanType) String() string

func (BooleanType) Type

func (BooleanType) Type() string

type BoundBooleanExprVisitor

type BoundBooleanExprVisitor[T any] interface {
	BooleanExprVisitor[T]

	VisitIn(BoundTerm, Set[Literal]) T
	VisitNotIn(BoundTerm, Set[Literal]) T
	VisitIsNan(BoundTerm) T
	VisitNotNan(BoundTerm) T
	VisitIsNull(BoundTerm) T
	VisitNotNull(BoundTerm) T
	VisitEqual(BoundTerm, Literal) T
	VisitNotEqual(BoundTerm, Literal) T
	VisitGreaterEqual(BoundTerm, Literal) T
	VisitGreater(BoundTerm, Literal) T
	VisitLessEqual(BoundTerm, Literal) T
	VisitLess(BoundTerm, Literal) T
	VisitStartsWith(BoundTerm, Literal) T
	VisitNotStartsWith(BoundTerm, Literal) T
}

BoundBooleanExprVisitor builds on BooleanExprVisitor by adding interface methods for visiting bound expressions, because we do casting of literals during binding you can assume that the BoundTerm and the Literal passed to a method have the same type.

type BoundLiteralPredicate

type BoundLiteralPredicate interface {
	BoundPredicate

	Literal() Literal
	AsUnbound(Reference, Literal) UnboundPredicate
}

BoundLiteralPredicate represents a bound boolean expression that utilizes a single literal as an argument, such as Equals or StartsWith.

type BoundPredicate

type BoundPredicate interface {
	BooleanExpression
	Ref() BoundReference
	Term() BoundTerm
}

BoundPredicate is a boolean predicate expression which has been bound to a schema. The underlying reference and term can be retrieved from it.

type BoundReference

type BoundReference interface {
	BoundTerm

	Field() NestedField
	Pos() int
	PosPath() []int
}

BoundReference is a named reference that has been bound to a particular field in a given schema.

type BoundSetPredicate

type BoundSetPredicate interface {
	BoundPredicate

	Literals() Set[Literal]
	AsUnbound(Reference, []Literal) UnboundPredicate
}

BoundSetPredicate is a bound expression that utilizes a set of literals such as In or NotIn

type BoundTerm

type BoundTerm interface {
	Term

	Equals(BoundTerm) bool
	Ref() BoundReference
	Type() Type
	// contains filtered or unexported methods
}

BoundTerm is a simple expression (typically a reference) that evaluates to a value and has been bound to a schema.

type BoundTransform

type BoundTransform struct {
	// contains filtered or unexported fields
}

func (*BoundTransform) Equals

func (b *BoundTransform) Equals(other BoundTerm) bool

func (*BoundTransform) Ref

func (b *BoundTransform) Ref() BoundReference

func (*BoundTransform) String

func (b *BoundTransform) String() string

func (*BoundTransform) Type

func (b *BoundTransform) Type() Type

type BoundUnaryPredicate

type BoundUnaryPredicate interface {
	BoundPredicate

	AsUnbound(Reference) UnboundPredicate
}

BoundUnaryPredicate is a bound predicate expression that has no arguments

type BucketTransform

type BucketTransform struct {
	NumBuckets int
}

BucketTransform transforms values into a bucket partition value. It is parameterized by a number of buckets. Bucket partition transforms use a 32-bit hash of the source value to produce a positive value by mod the bucket number.

func (BucketTransform) Apply

func (t BucketTransform) Apply(value Optional[Literal]) Optional[Literal]

func (BucketTransform) CanTransform added in v0.4.0

func (BucketTransform) CanTransform(t Type) bool

func (BucketTransform) Equals

func (t BucketTransform) Equals(other Transform) bool

func (BucketTransform) MarshalText

func (t BucketTransform) MarshalText() ([]byte, error)

func (BucketTransform) PreservesOrder added in v0.2.0

func (BucketTransform) PreservesOrder() bool

func (BucketTransform) Project

func (t BucketTransform) Project(name string, pred BoundPredicate) (UnboundPredicate, error)

func (BucketTransform) ResultType

func (BucketTransform) ResultType(Type) Type

func (BucketTransform) String

func (t BucketTransform) String() string

func (BucketTransform) ToHumanStr added in v0.2.0

func (BucketTransform) ToHumanStr(val any) string

func (BucketTransform) Transformer

func (t BucketTransform) Transformer(src Type) func(any) Optional[int32]

type Comparator

type Comparator[T LiteralType] func(v1, v2 T) int

Comparator is a comparison function for specific literal types:

returns 0 if v1 == v2
returns <0 if v1 < v2
returns >0 if v1 > v2

type DataFile

type DataFile interface {
	// ContentType is the type of the content stored by the data file,
	// either Data, Equality deletes, or Position deletes. All v1 files
	// are Data files.
	ContentType() ManifestEntryContent
	// FilePath is the full URI for the file, complete with FS scheme.
	FilePath() string
	// FileFormat is the format of the data file, AVRO, Orc, or Parquet.
	FileFormat() FileFormat
	// Partition returns a mapping of field id to partition value for
	// each of the partition spec's fields.
	Partition() map[int]any
	// Count returns the number of records in this file.
	Count() int64
	// FileSizeBytes is the total file size in bytes.
	FileSizeBytes() int64
	// ColumnSizes is a mapping from column id to the total size on disk
	// of all regions that store the column. Does not include bytes
	// necessary to read other columns, like footers. Map will be nil for
	// row-oriented formats (avro).
	ColumnSizes() map[int]int64
	// ValueCounts is a mapping from column id to the number of values
	// in the column, including null and NaN values.
	ValueCounts() map[int]int64
	// NullValueCounts is a mapping from column id to the number of
	// null values in the column.
	NullValueCounts() map[int]int64
	// NaNValueCounts is a mapping from column id to the number of NaN
	// values in the column.
	NaNValueCounts() map[int]int64
	// DistictValueCounts is a mapping from column id to the number of
	// distinct values in the column. Distinct counts must be derived
	// using values in the file by counting or using sketches, but not
	// using methods like merging existing distinct counts.
	DistinctValueCounts() map[int]int64
	// LowerBoundValues is a mapping from column id to the lower bounded
	// value of the column, serialized as binary. Each value in the column
	// must be less than or requal to all non-null, non-NaN values in the
	// column for the file.
	LowerBoundValues() map[int][]byte
	// UpperBoundValues is a mapping from column id to the upper bounded
	// value of the column, serialized as binary. Each value in the column
	// must be greater than or equal to all non-null, non-NaN values in
	// the column for the file.
	UpperBoundValues() map[int][]byte
	// KeyMetadata is implementation-specific key metadata for encryption.
	KeyMetadata() []byte
	// SplitOffsets are the split offsets for the data file. For example,
	// all row group offsets in a Parquet file. Must be sorted ascending.
	SplitOffsets() []int64
	// EqualityFieldIDs are used to determine row equality in equality
	// delete files. It is required when the content type is
	// EntryContentEqDeletes.
	EqualityFieldIDs() []int
	// SortOrderID returns the id representing the sort order for this
	// file, or nil if there is no sort order.
	SortOrderID() *int
	// SpecID returns the partition spec id for this data file, inherited
	// from the manifest that the data file was read from
	SpecID() int32
	// FirstRowID returns the first row ID for this data file ( v3+ only )
	FirstRowID() *int64
	// ReferencedDataFile returns the location of the data file that deletion vector reference
	ReferencedDataFile() *string
	// ContentOffset returns the offset in the file where the content starts ( v3+ only )
	ContentOffset() *int64
	// ContentSizeInBytes returns the length of referenced contented stored in the file (v3+ only)
	ContentSizeInBytes() *int64
}

DataFile is the interface for reading the information about a given data file indicated by an entry in a manifest list.

type DataFileBuilder

type DataFileBuilder struct {
	// contains filtered or unexported fields
}

DataFileBuilder is a helper for building a data file struct which will conform to the DataFile interface.

func NewDataFileBuilder

func NewDataFileBuilder(
	spec PartitionSpec,
	content ManifestEntryContent,
	path string,
	format FileFormat,
	fieldIDToPartitionData map[int]any,
	fieldIDToLogicalType map[int]string,
	fieldIDToFixedSize map[int]int,
	recordCount int64,
	fileSize int64,
) (*DataFileBuilder, error)

NewDataFileBuilder is passed all of the required fields and then allows all of the optional fields to be set by calling the corresponding methods before calling DataFileBuilder.Build to construct the object.

func (*DataFileBuilder) BlockSizeInBytes

func (b *DataFileBuilder) BlockSizeInBytes(size int64) *DataFileBuilder

BlockSizeInBytes sets the block size in bytes for the data file. Deprecated in v2.

func (*DataFileBuilder) Build

func (b *DataFileBuilder) Build() DataFile

func (*DataFileBuilder) ColumnSizes

func (b *DataFileBuilder) ColumnSizes(sizes map[int]int64) *DataFileBuilder

ColumnSizes sets the column sizes for the data file.

func (*DataFileBuilder) ContentOffset added in v0.4.0

func (b *DataFileBuilder) ContentOffset(offset int64) *DataFileBuilder

func (*DataFileBuilder) ContentSizeInBytes added in v0.4.0

func (b *DataFileBuilder) ContentSizeInBytes(size int64) *DataFileBuilder

func (*DataFileBuilder) DistinctValueCounts deprecated

func (b *DataFileBuilder) DistinctValueCounts(counts map[int]int64) *DataFileBuilder

DistinctValueCounts sets the distinct value counts for the data file.

Deprecated: distinct_counts (field 111) is deprecated in every version of the Iceberg spec (apache/iceberg#12182). The Avro manifest-entry schemas omit the field for v1, v2, and v3, so values set here are not transported in manifests written by this library. The setter is retained for round-tripping legacy DataFiles read from older manifests; new code should not call it.

func (*DataFileBuilder) EqualityFieldIDs

func (b *DataFileBuilder) EqualityFieldIDs(ids []int) *DataFileBuilder

EqualityFieldIDs sets the equality field ids for the data file.

func (*DataFileBuilder) FirstRowID added in v0.4.0

func (b *DataFileBuilder) FirstRowID(id int64) *DataFileBuilder

func (*DataFileBuilder) KeyMetadata

func (b *DataFileBuilder) KeyMetadata(key []byte) *DataFileBuilder

KeyMetadata sets the key metadata for the data file.

func (*DataFileBuilder) LowerBoundValues

func (b *DataFileBuilder) LowerBoundValues(bounds map[int][]byte) *DataFileBuilder

LowerBoundValues sets the lower bound values for the data file.

func (*DataFileBuilder) NaNValueCounts

func (b *DataFileBuilder) NaNValueCounts(counts map[int]int64) *DataFileBuilder

NaNValueCounts sets the NaN value counts for the data file.

func (*DataFileBuilder) NullValueCounts

func (b *DataFileBuilder) NullValueCounts(counts map[int]int64) *DataFileBuilder

NullValueCounts sets the null value counts for the data file.

func (*DataFileBuilder) ReferencedDataFile added in v0.4.0

func (b *DataFileBuilder) ReferencedDataFile(path string) *DataFileBuilder

func (*DataFileBuilder) SortOrderID

func (b *DataFileBuilder) SortOrderID(id int) *DataFileBuilder

SortOrderID sets the sort order id for the data file.

func (*DataFileBuilder) SplitOffsets

func (b *DataFileBuilder) SplitOffsets(offsets []int64) *DataFileBuilder

SplitOffsets sets the split offsets for the data file.

func (*DataFileBuilder) UpperBoundValues

func (b *DataFileBuilder) UpperBoundValues(bounds map[int][]byte) *DataFileBuilder

UpperBoundValues sets the upper bound values for the data file.

func (*DataFileBuilder) ValueCounts

func (b *DataFileBuilder) ValueCounts(counts map[int]int64) *DataFileBuilder

ValueCounts sets the value counts for the data file.

type Date

type Date int32

func (Date) ToTime

func (d Date) ToTime() time.Time

type DateLiteral

type DateLiteral Date

func (DateLiteral) Any added in v0.2.0

func (d DateLiteral) Any() any

func (DateLiteral) Comparator

func (DateLiteral) Comparator() Comparator[Date]

func (DateLiteral) Decrement

func (d DateLiteral) Decrement() Literal

func (DateLiteral) Equals

func (d DateLiteral) Equals(other Literal) bool

func (DateLiteral) Increment

func (d DateLiteral) Increment() Literal

func (DateLiteral) MarshalBinary

func (d DateLiteral) MarshalBinary() (data []byte, err error)

func (DateLiteral) String

func (d DateLiteral) String() string

func (DateLiteral) To

func (d DateLiteral) To(t Type) (Literal, error)

func (DateLiteral) Type

func (d DateLiteral) Type() Type

func (*DateLiteral) UnmarshalBinary

func (d *DateLiteral) UnmarshalBinary(data []byte) error

func (DateLiteral) Value

func (d DateLiteral) Value() Date

type DateType

type DateType struct{}

DateType represents a calendar date without a timezone or time, represented as a 32-bit integer denoting the number of days since the unix epoch.

func (DateType) Equals

func (DateType) Equals(other Type) bool

func (DateType) String

func (DateType) String() string

func (DateType) Type

func (DateType) Type() string

type DayTransform

type DayTransform struct{}

DayTransform transforms a datetime value into a date value.

func (DayTransform) Apply

func (DayTransform) Apply(value Optional[Literal]) (out Optional[Literal])

func (DayTransform) CanTransform added in v0.4.0

func (t DayTransform) CanTransform(sourceType Type) bool

func (DayTransform) Equals

func (DayTransform) Equals(other Transform) bool

func (DayTransform) MarshalText

func (t DayTransform) MarshalText() ([]byte, error)

func (DayTransform) PreservesOrder added in v0.2.0

func (DayTransform) PreservesOrder() bool

func (DayTransform) Project

func (t DayTransform) Project(name string, pred BoundPredicate) (UnboundPredicate, error)

func (DayTransform) ResultType

func (DayTransform) ResultType(Type) Type

func (DayTransform) String

func (DayTransform) String() string

func (DayTransform) ToHumanStr added in v0.2.0

func (DayTransform) ToHumanStr(val any) string

func (DayTransform) Transformer

func (DayTransform) Transformer(src Type) (func(any) Optional[int32], error)

type Decimal

type Decimal struct {
	Val   decimal.Decimal128
	Scale int
}

func (Decimal) String added in v0.2.0

func (d Decimal) String() string

type DecimalLiteral

type DecimalLiteral Decimal

func (DecimalLiteral) Any added in v0.2.0

func (d DecimalLiteral) Any() any

func (DecimalLiteral) Comparator

func (DecimalLiteral) Comparator() Comparator[Decimal]

func (DecimalLiteral) Decrement

func (d DecimalLiteral) Decrement() Literal

func (DecimalLiteral) Equals

func (d DecimalLiteral) Equals(other Literal) bool

func (DecimalLiteral) Increment

func (d DecimalLiteral) Increment() Literal

func (DecimalLiteral) MarshalBinary

func (d DecimalLiteral) MarshalBinary() (data []byte, err error)

func (DecimalLiteral) String

func (d DecimalLiteral) String() string

func (DecimalLiteral) To

func (d DecimalLiteral) To(t Type) (Literal, error)

func (DecimalLiteral) Type

func (d DecimalLiteral) Type() Type

func (*DecimalLiteral) UnmarshalBinary

func (d *DecimalLiteral) UnmarshalBinary(data []byte) error

func (DecimalLiteral) Value

func (d DecimalLiteral) Value() Decimal

type DecimalType

type DecimalType struct {
	// contains filtered or unexported fields
}

func DecimalTypeOf

func DecimalTypeOf(prec, scale int) DecimalType

func (DecimalType) Equals

func (d DecimalType) Equals(other Type) bool

func (DecimalType) Precision

func (d DecimalType) Precision() int

func (DecimalType) Scale

func (d DecimalType) Scale() int

func (DecimalType) String

func (d DecimalType) String() string

func (DecimalType) Type

func (d DecimalType) Type() string

type FieldSummary

type FieldSummary struct {
	ContainsNull bool    `avro:"contains_null"`
	ContainsNaN  *bool   `avro:"contains_nan"`
	LowerBound   *[]byte `avro:"lower_bound"`
	UpperBound   *[]byte `avro:"upper_bound"`
}

type FileFormat

type FileFormat string

FileFormat defines constants for the format of data files.

const (
	AvroFile    FileFormat = "AVRO"
	OrcFile     FileFormat = "ORC"
	ParquetFile FileFormat = "PARQUET"
	PuffinFile  FileFormat = "PUFFIN"
)

func FileFormatFromString added in v0.6.0

func FileFormatFromString(s string) (FileFormat, error)

FileFormatFromString parses a file format string (case-insensitive).

type FixedLiteral

type FixedLiteral []byte

func (FixedLiteral) Any added in v0.2.0

func (f FixedLiteral) Any() any

func (FixedLiteral) Comparator

func (FixedLiteral) Comparator() Comparator[[]byte]

func (FixedLiteral) Equals

func (f FixedLiteral) Equals(other Literal) bool

func (FixedLiteral) MarshalBinary

func (f FixedLiteral) MarshalBinary() (data []byte, err error)

func (FixedLiteral) String

func (f FixedLiteral) String() string

func (FixedLiteral) To

func (f FixedLiteral) To(typ Type) (Literal, error)

func (FixedLiteral) Type

func (f FixedLiteral) Type() Type

func (*FixedLiteral) UnmarshalBinary

func (f *FixedLiteral) UnmarshalBinary(data []byte) error

func (FixedLiteral) Value

func (f FixedLiteral) Value() []byte

type FixedType

type FixedType struct {
	// contains filtered or unexported fields
}

func FixedTypeOf

func FixedTypeOf(n int) FixedType

func (FixedType) Equals

func (f FixedType) Equals(other Type) bool

func (FixedType) Len

func (f FixedType) Len() int

func (FixedType) String

func (f FixedType) String() string

func (FixedType) Type

func (f FixedType) Type() string

type Float32Literal

type Float32Literal float32

func (Float32Literal) Any added in v0.2.0

func (f Float32Literal) Any() any

func (Float32Literal) Comparator

func (Float32Literal) Comparator() Comparator[float32]

func (Float32Literal) Equals

func (f Float32Literal) Equals(other Literal) bool

func (Float32Literal) MarshalBinary

func (f Float32Literal) MarshalBinary() (data []byte, err error)

func (Float32Literal) String

func (f Float32Literal) String() string

func (Float32Literal) To

func (f Float32Literal) To(t Type) (Literal, error)

func (Float32Literal) Type

func (f Float32Literal) Type() Type

func (*Float32Literal) UnmarshalBinary

func (f *Float32Literal) UnmarshalBinary(data []byte) error

func (Float32Literal) Value

func (f Float32Literal) Value() float32

type Float32Type

type Float32Type struct{}

Float32Type is the "float" type in the iceberg spec.

func (Float32Type) Equals

func (Float32Type) Equals(other Type) bool

func (Float32Type) String

func (Float32Type) String() string

func (Float32Type) Type

func (Float32Type) Type() string

type Float64Literal

type Float64Literal float64

func (Float64Literal) Any added in v0.2.0

func (f Float64Literal) Any() any

func (Float64Literal) Comparator

func (Float64Literal) Comparator() Comparator[float64]

func (Float64Literal) Equals

func (f Float64Literal) Equals(other Literal) bool

func (Float64Literal) MarshalBinary

func (f Float64Literal) MarshalBinary() (data []byte, err error)

func (Float64Literal) String

func (f Float64Literal) String() string

func (Float64Literal) To

func (f Float64Literal) To(t Type) (Literal, error)

func (Float64Literal) Type

func (f Float64Literal) Type() Type

func (*Float64Literal) UnmarshalBinary

func (f *Float64Literal) UnmarshalBinary(data []byte) error

func (Float64Literal) Value

func (f Float64Literal) Value() float64

type Float64Type

type Float64Type struct{}

Float64Type represents the "double" type of the iceberg spec.

func (Float64Type) Equals

func (Float64Type) Equals(other Type) bool

func (Float64Type) String

func (Float64Type) String() string

func (Float64Type) Type

func (Float64Type) Type() string

type HourTransform

type HourTransform struct{}

HourTransform transforms a datetime value into an hour value.

func (HourTransform) Apply

func (HourTransform) Apply(value Optional[Literal]) (out Optional[Literal])

func (HourTransform) CanTransform added in v0.4.0

func (t HourTransform) CanTransform(sourceType Type) bool

func (HourTransform) Equals

func (HourTransform) Equals(other Transform) bool

func (HourTransform) MarshalText

func (t HourTransform) MarshalText() ([]byte, error)

func (HourTransform) PreservesOrder added in v0.2.0

func (HourTransform) PreservesOrder() bool

func (HourTransform) Project

func (t HourTransform) Project(name string, pred BoundPredicate) (UnboundPredicate, error)

func (HourTransform) ResultType

func (HourTransform) ResultType(Type) Type

func (HourTransform) String

func (HourTransform) String() string

func (HourTransform) ToHumanStr added in v0.2.0

func (HourTransform) ToHumanStr(val any) string

func (HourTransform) Transformer

func (HourTransform) Transformer(src Type) (func(any) Optional[int32], error)

type IdentityTransform

type IdentityTransform struct{}

IdentityTransform uses the identity function, performing no transformation but instead partitioning on the value itself.

func (IdentityTransform) Apply

func (IdentityTransform) CanTransform added in v0.4.0

func (IdentityTransform) CanTransform(t Type) bool

func (IdentityTransform) Equals

func (IdentityTransform) Equals(other Transform) bool

func (IdentityTransform) MarshalText

func (t IdentityTransform) MarshalText() ([]byte, error)

func (IdentityTransform) PreservesOrder added in v0.2.0

func (IdentityTransform) PreservesOrder() bool

func (IdentityTransform) Project

func (IdentityTransform) ResultType

func (IdentityTransform) ResultType(t Type) Type

func (IdentityTransform) String

func (IdentityTransform) String() string

func (IdentityTransform) ToHumanStr added in v0.2.0

func (IdentityTransform) ToHumanStr(val any) string

type Int32Literal

type Int32Literal int32

func (Int32Literal) Any added in v0.2.0

func (i Int32Literal) Any() any

func (Int32Literal) Comparator

func (Int32Literal) Comparator() Comparator[int32]

func (Int32Literal) Decrement

func (i Int32Literal) Decrement() Literal

func (Int32Literal) Equals

func (i Int32Literal) Equals(other Literal) bool

func (Int32Literal) Increment

func (i Int32Literal) Increment() Literal

func (Int32Literal) MarshalBinary

func (i Int32Literal) MarshalBinary() (data []byte, err error)

func (Int32Literal) String

func (i Int32Literal) String() string

func (Int32Literal) To

func (i Int32Literal) To(t Type) (Literal, error)

func (Int32Literal) Type

func (i Int32Literal) Type() Type

func (*Int32Literal) UnmarshalBinary

func (i *Int32Literal) UnmarshalBinary(data []byte) error

func (Int32Literal) Value

func (i Int32Literal) Value() int32

type Int32Type

type Int32Type struct{}

Int32Type is the "int"/"integer" type of the iceberg spec.

func (Int32Type) Equals

func (Int32Type) Equals(other Type) bool

func (Int32Type) String

func (Int32Type) String() string

func (Int32Type) Type

func (Int32Type) Type() string

type Int64Literal

type Int64Literal int64

func (Int64Literal) Any added in v0.2.0

func (i Int64Literal) Any() any

func (Int64Literal) Comparator

func (Int64Literal) Comparator() Comparator[int64]

func (Int64Literal) Decrement

func (i Int64Literal) Decrement() Literal

func (Int64Literal) Equals

func (i Int64Literal) Equals(other Literal) bool

func (Int64Literal) Increment

func (i Int64Literal) Increment() Literal

func (Int64Literal) MarshalBinary

func (i Int64Literal) MarshalBinary() (data []byte, err error)

func (Int64Literal) String

func (i Int64Literal) String() string

func (Int64Literal) To