Tuesday, December 19, 2017

Hadoop Real Time interview questions

BigData Realtime Interview Questions


  1. which table is faster when we are querying in parquet and textformat table in hive..?
  2.  what are the narrow transformations and wide range transformations in spark ..??
  3. what are the types of streaming sources in spark streaming ..??
               i) Basic sources  --reading from File , socket connections..etc(Available in spark core API)
               ii)Advanced sources --kafkaUtils, TwitterUtils,Flume..etc

     4. what are the components of hive OR explain the execution of Hive query ..??
     5. what is lineage graph in sark ..??
     6. Explain Map-side join in Hive ..??
     7. Explain difference between clustered by key and distributed by in Hive ..??
     8. what is the  default partitioner in Mapreduce ..??
     9. what is the Difference between reduce and reduceByKey in Spark ..?
    10. what is the difference between client mode and cluster mode execution in spark ..??
    11. what are the transformations and actions in spark ..??
    12 . what is accumulators in spark ..??
    13. what is the role of akka framework in spark .??
    14. Difference between persist and cache in spark ..??
    15. types of persist in spark ..??
    16. Explain DAG in spark ..??
    17. Difference between Data Frame and RDD ..??
    18. Difference between Dataset and Data Frame ..??
    19. How to fetch last element in RDD ..??
    20.  How to create schema for a text file in spark ..??
    21. How spark knows it's a streaming job ..??
    22. what are the default no.of mappers in sqoop ..??
    23. How to fetch messages from particular offset to offset in kafka ..??
    24. what is Synchronous and Asynchronous in kafka ..??
    25. what is commit in kafka ..??
    26.  what is .index file in kafka ..??

             

No comments:

Post a Comment