spark-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Egor Pahomov (JIRA)" <>
Subject [jira] [Commented] (SPARK-19524) newFilesOnly does not work according to docs.
Date Thu, 09 Feb 2017 19:15:41 GMT


Egor Pahomov commented on SPARK-19524:

I'm really confused. I expected "new" to be the files created after start of streaming job
and old ones to be everything else in the folder. If we change definition of "new", than I
believe everything consistent between each other. It's just I'm not sure that this "new" definition
is very intuitive. I want to process everything in folder - existing and upcoming. I use this
flag. And now it turns out, that this flag has it's own definition of "new". My be I'm not
correct to call it a bug, but isn't it all very confusing for person, who does not really
know who everything works inside? 

> newFilesOnly does not work according to docs. 
> ----------------------------------------------
>                 Key: SPARK-19524
>                 URL:
>             Project: Spark
>          Issue Type: Bug
>          Components: DStreams
>    Affects Versions: 2.0.2
>            Reporter: Egor Pahomov
> Docs says:
> newFilesOnly
> Should process only new files and ignore existing files in the directory
> It's not working. 
says, that it shouldn't work as expected. 
not clear at all in terms, what code tries to do

This message was sent by Atlassian JIRA

To unsubscribe, e-mail:
For additional commands, e-mail:

View raw message