hive-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Chinmay Kulkarni (Jira)" <>
Subject [jira] [Commented] (HIVE-23851) MSCK REPAIR Command With Partition Filtering Fails While Dropping Partitions
Date Tue, 27 Jul 2021 18:11:00 GMT


Chinmay Kulkarni commented on HIVE-23851:

[~ashutoshc] [~srahman] [~kgyrtkirk] [~prasanth_j] I am trying to implement an API in a downstream
project (Trino) which would eventually call the Hive {{ThriftHiveMetastore.Iface#drop_partitions_req()}}

I see that in this Jira, we have a separate implementation of the {{PartitionExpressionProxy}}
class which doesn't depend on any hive-exec classes and is accessible to a standalone HMS.
I wish to do the same thing.

Is there any way to support calling bulkdropPartitions APIs for  downstream projects without
having to make changes to Hive and implement their own {{PartitionExpressionProxy}} class?
Any discussions/examples of projects that may have done this?

> MSCK REPAIR Command With Partition Filtering Fails While Dropping Partitions
> ----------------------------------------------------------------------------
>                 Key: HIVE-23851
>                 URL:
>             Project: Hive
>          Issue Type: Bug
>    Affects Versions: 4.0.0
>            Reporter: Syed Shameerur Rahman
>            Assignee: Syed Shameerur Rahman
>            Priority: Major
>              Labels: pull-request-available
>             Fix For: 4.0.0
>          Time Spent: 5.5h
>  Remaining Estimate: 0h
> *Steps to reproduce:*
> # Create external table
> # Run msck command to sync all the partitions with metastore
> # Remove one of the partition path
> # Run msck repair with partition filtering
> *Stack Trace:*
> {code:java}
>  2020-07-15T02:10:29,045 ERROR [4dad298b-28b1-4e6b-94b6-aa785b60c576 main] ppr.PartitionExpressionForMetastore:
Failed to deserialize the expression
>  java.lang.IndexOutOfBoundsException: Index: 110, Size: 0
>  at java.util.ArrayList.rangeCheck( ~[?:1.8.0_192]
>  at java.util.ArrayList.get( ~[?:1.8.0_192]
>  at
>  at
>  at ~[hive-exec-4.0.0-SNAPSHOT.jar:4.0.0-SNAPSHOT]
>  at org.apache.hadoop.hive.ql.exec.SerializationUtilities$KryoWithHooks.readObject(
>  at org.apache.hadoop.hive.ql.exec.SerializationUtilities.deserializeObjectFromKryo(
>  at org.apache.hadoop.hive.ql.exec.SerializationUtilities.deserializeExpressionFromKryo(
>  at org.apache.hadoop.hive.ql.optimizer.ppr.PartitionExpressionForMetastore.deserializeExpr(
>  at org.apache.hadoop.hive.ql.optimizer.ppr.PartitionExpressionForMetastore.convertExprToFilter(
>  at org.apache.hadoop.hive.metastore.PartFilterExprUtil.makeExpressionTree(
>  at org.apache.hadoop.hive.metastore.ObjectStore.getPartitionsByExprInternal(
>  at org.apache.hadoop.hive.metastore.VerifyingObjectStore.getPartitionsByExpr(
>  at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method) ~[?:1.8.0_192]
>  at sun.reflect.NativeMethodAccessorImpl.invoke( ~[?:1.8.0_192]
> {code}
> *Cause:*
> In case of msck repair with partition filtering we expect expression proxy class to be
set as PartitionExpressionForMetastore (
), While dropping partition we serialize the drop partition filter expression as (
) which is incompatible during deserializtion happening in PartitionExpressionForMetastore
) hence the query fails with Failed to deserialize the expression.
> *Solutions*:
> I could think of two approaches to this problem
> # Since PartitionExpressionForMetastore is required only during parition pruning step,
We can switch back the expression proxy class to MsckPartitionExpressionProxy once the partition
pruning step is done.
> # The other solution is to make serialization process in msck drop partition filter expression
compatible with the one with PartitionExpressionForMetastore, We can do this via Reflection
since the drop partition serialization happens in Msck class (standadlone-metatsore) by this
way we can completely remove the need for class MsckPartitionExpressionProxy and this also
helps to reduce the complexity of Msck Repair command with parition filtering to work with
ease (no need to set the expression proxyClass config).
> I am personally inclined to the 2nd approach. Before moving on i want to know if this
is the best approach or is there any other better/easier approach to solve this problem.
> PS: qtest added in HIVE-22957 mainly focused on adding missing partition. Forgot to add
case for dropping partition.

This message was sent by Atlassian Jira

View raw message